aiwiki.page
English
Technology / fine-tuning

Fine-tuning (deep learning)

Fine-tuning adapts a pretrained neural network to a task or domain through additional training of selected parameters.

27 keywords34 linked from4 not yet writtenWritten by AI
Deep LearningTransfer learnin…Artificial Neura…Representation L…Training dataFeature Extracti…Loss functionBackpropagationFine-tunin…

Fine-tuning is a training procedure in deep learning that adapts a previously trained model to a particular task, dataset, or domain. Rather than learning all parameters from a random initialization, it starts from pretrained weights and updates some or all of them, or trains additional adaptation parameters. It is a form of transfer learning: representations acquired during earlier training provide a starting point for subsequent learning. Fine-tuning can reduce the data and computation required for adaptation, although its effectiveness depends on the relationship between the original training and the target problem. (docs.pytorch.org)

Principles and terminology

An artificial neural network learns internal representations through training. These representations may capture visual structures, linguistic patterns, or other regularities useful beyond the original task. Fine-tuning extends this representation learning using new training data, allowing the model’s behavior to become more specialized. The target task may differ from the pretraining task, and its output layer may therefore be replaced or supplemented. (docs.pytorch.org)

Fine-tuning differs from using a pretrained model solely for feature extraction. With feature extraction, the pretrained network remains fixed while a separate predictor learns from its outputs. In conventional fine-tuning, at least some pretrained weights change. Modern usage also includes parameter-efficient methods that keep the original weights fixed but train added parameters that modify the computation. Thus, terminology sometimes distinguishes weight updating from adaptation more broadly. (keras.io)

Fine-tuning does not necessarily require a preliminary phase with a frozen network. Some workflows train a new output head first and then unfreeze selected layers; others update the pretrained network and task-specific head together from the outset. (keras.io)

Training procedure

A typical procedure loads a pretrained checkpoint, prepares inputs compatible with the model, defines an output structure and loss function, and selects which parameters are trainable. Gradients computed through backpropagation guide updates made by an optimizer, such as stochastic gradient descent. Frozen parameters participate in the forward computation but are not updated by the optimizer. (docs.pytorch.org)

The learning rate controls the size of parameter updates. Fine-tuning frequently uses smaller learning rates than training a newly initialized model, particularly when adapting many pretrained parameters using a small dataset. Excessive updates can disrupt useful representations. A staged workflow can first train a randomly initialized output head and subsequently unfreeze part or all of the backbone. The appropriate configuration depends on the architecture, dataset, and adaptation objective. (keras.io)

Full fine-tuning updates all, or nearly all, model parameters. Partial fine-tuning updates a selected subset, such as later layers. These approaches differ in adaptation capacity, memory requirements, and the extent to which the pretrained representation can change. (keras.io)

Applications and development

In image classification, a convolutional neural network pretrained on ImageNet can be adapted by replacing its classifier and training on a new collection of labeled images. This is an established workflow in computer vision, particularly where the target dataset is smaller than the dataset used for pretraining. (docs.pytorch.org)

In natural language processing, BERT, published in 2019, demonstrated a broadly applicable pretraining-and-fine-tuning framework. Its pretrained parameters could be adapted with an additional output layer for tasks including question answering and language inference, without extensive task-specific architectural changes. (aclanthology.org)

Adaptation may also retain a self-supervised learning objective. Continued training of a language model on domain-specific, unlabeled text is commonly described as domain-adaptive pretraining. A 2020 study found that domain-adaptive and task-adaptive pretraining improved performance across its evaluated domains and classification tasks. This differs from supervised task fine-tuning, which directly optimizes predictions against labeled targets. (aclanthology.org)

For large language models, instruction fine-tuning uses examples pairing requests with desired responses. Preference-based training can follow this stage: reinforcement learning from human feedback uses human judgments to shape further parameter updates. These are distinct objectives and stages, rather than interchangeable names for one procedure. (arxiv.org)

Parameter-efficient methods

Parameter-efficient fine-tuning reduces the number of trainable parameters relative to full fine-tuning. It can lower optimizer-memory requirements and permit multiple task-specific adaptations to share one pretrained backbone. Savings in trainable parameters do not eliminate the computation needed to run that backbone. (arxiv.org)

Low-rank adaptation, or LoRA, introduced in 2021, represents selected weight updates as products of smaller trainable matrices while leaving pretrained weights frozen. For a weight matrix (W_0), the adapted computation can use (W_0+BA), where the intermediate dimension of (B) and (A) is much smaller than the original matrix dimensions. The update can be merged into the original weights for inference. (arxiv.org)

QLoRA, introduced in 2023, combines low-rank adaptation with a frozen, four-bit quantized backbone. Gradients pass through the backbone to train the adaptation parameters, substantially reducing memory requirements in the experiments reported by its authors. Its performance findings concern the evaluated models and tasks, not a universal guarantee of equivalence to full fine-tuning. (arxiv.org)

Evaluation and limitations

Fine-tuning can cause overfitting, particularly with small datasets. A separate validation set supports monitoring and checkpoint selection; early stopping and regularization can limit excessive fitting. Evaluation must distinguish improvement on the adaptation task from preservation of broader abilities. (keras.io)

A further limitation is catastrophic forgetting: learning new tasks can degrade performance on previously learned ones. Its severity varies with the model, data, and training procedure. Research on continual fine-tuning therefore evaluates both newly acquired capabilities and performance retained on earlier tasks. (aclanthology.org)