Transfer Learning Explained: Reusing and Fine-Tuning a Model

What is transfer learning

Transfer learning reuses knowledge learned from one dataset or task to help with another. For a neural network, this can mean keeping a pretrained feature extractor fixed and training a new output head, or fine-tuning some or all of the pretrained parameters. The adapted model still needs evaluation on relevant unseen data.

A pretrained model has already learned parameters, such as weights and biases, through earlier training. These values shape how it transforms inputs into outputs. They can provide a useful starting point for a new problem. If those terms are unfamiliar, Neural Networks Explained introduces how layers and their adjustable values work together.

The source task is the original learning problem. The target task is the problem you want to solve now. A domain describes the kinds of inputs and how they are distributed: studio product photographs and outdoor phone photographs can belong to different domains even when both contain shoes.

The task, domain or both may change. Transfer learning is a reuse strategy, and fine-tuning is one way to carry it out. It can work alongside the approaches covered in Types of Machine Learning. Here, we focus on neural networks within deep learning, using one product-image example.

Three ways to build the target model

Train a new head with fixed features

Divide the model into a backbone and a head. The backbone, or feature extractor, transforms an image into a numerical representation. The head turns that representation into task-specific outputs.

Keep the pretrained backbone and replace the original head with a new one initialized with random weights. Freeze the backbone’s parameters so the optimizer, which applies training updates, leaves them unchanged. Train the new head on the target labels. The PyTorch transfer-learning tutorial demonstrates this fixed-feature approach alongside fine-tuning.

The backbone still processes images. Its parameters remain fixed, but different images produce different features.

For example, the same feature extractor can produce different numerical descriptions of a shoe and a bag. The head learns how those descriptions relate to the target labels.

Fine-tune pretrained parameters

Fine-tuning allows selected pretrained parameters to change during target training. You might unfreeze only the later backbone layers, or train the whole backbone together with the new head.

This lets the representation adapt to the target photos. Parameters in any remaining frozen layers stay fixed. A learning rate controls the scale of training updates; a smaller rate is a common precaution when adjusting useful pretrained features.

There is one implementation wrinkle: frozen parameters do not automatically freeze every stored value. Batch-normalization layers can maintain running statistics outside optimizer updates. Training mode and those statistics need deliberate handling, as the Keras transfer-learning guide explains.

Train from scratch

A scratch model starts without learned source weights. In our comparison, a comparable backbone and a new head start with random weights and train on the target data.

This is the no-transfer alternative. It helps answer whether pretraining actually improves the outcome. Give each approach appropriate training settings and compare both quality and practical cost. Making every run equally short can disadvantage a method that needs longer to learn.

Fixed-feature transfer trains the head, fine-tuning also updates pretrained layers, and scratch training learns both parts.
The starting weights and the parts allowed to train distinguish these approaches. View the diagram at full size.

A product image example you can follow

Define the source and target

Return to the shoes, bags and hats in Convolutional Neural Networks Explained. This time, imagine a shop wants to classify new product photographs.

The source model is a convolutional neural network (CNN) trained to classify general-object photographs using its original labels. We reuse its visual backbone. The target dataset contains shop photos, each showing one product labeled shoes, bags or hats.

Replace the source classifier with a head that produces three scores, one per target class. The original classes need not match these three. Replacing the head changes the model’s structure; training learns the new head’s parameters. Merely renaming the source model’s output labels would not teach it this new classification task.

We are transferring a learned way of representing images, rather than looking up stored source answers. The representation may capture patterns useful for the shop, but usefulness is something to test. Image Classification Explained covers the relationship between an image, a class and its score.

A general-object source task is adapted to classifying shoes, bags and hats in shop photos.
The target can change the task, the image domain, or both. View the diagram at full size.

Follow one shoe image through training

Take a training photograph labeled shoes. First, prepare its pixels in the format expected by the backbone. The backbone computes features, and the new head combines them into scores for shoes, bags and hats.

Suppose the head currently favors bags. A loss function measures the mismatch between the output and the shoes label. Backpropagation calculates gradients, which describe how parameter changes affect the loss. An optimizer uses them to update the parameters allowed to train.

With fixed features, only the head receives optimizer updates. With fine-tuning, the head and selected pretrained backbone parameters receive updates. From scratch, both parts learn without the source model’s learned starting point.

One update does not guarantee a correct prediction. Training repeats across examples, while validation checks whether the resulting model handles held-out photos better.

After training, ordinary inference processes a new photograph using the selected model without updating its learned parameters. The fuller prediction, loss and update loop is covered in How Deep Learning Works.

A shoe label guides head updates and, during fine-tuning, updates to selected pretrained layers.
The labeled example guides training; its output is not a copied source answer. View the diagram at full size.

How to adapt and evaluate a pretrained model

Prepare data and a baseline

Define the labels, intended photo conditions and success measure. If missing hats matters especially, check hat errors rather than relying only on overall accuracy.

Collect representative examples. If the intended use includes customer phone photos, include that variety rather than evaluating only on studio images. Keep near-duplicates and different views of the same product in one split, so familiar products cannot leak across training and evaluation. Split before creating augmented copies, such as crops, or fitting preprocessing to target data.

Choose a suitable pretrained backbone and check its license and documented input preprocessing. Fit any additional learned preprocessing on training data only. The scikit-learn guidance on data leakage explains why held-out information must stay out of fitting.

Set up a comparable scratch baseline using the same target splits and evaluation measures. Record the available data, training budget and hardware.

Train and validate

Start by training the new head with the backbone fixed. Use validation data to choose settings and when to stop. Then, if useful, try unfreezing selected layers and fine-tuning with a smaller learning rate. This staged workflow is described in the Keras guide.

Compare the fine-tuned result with both the head-only model and the scratch baseline. Choose using validation evidence and cost, allowing sensible tuning for each approach. Keep the test set outside these decisions.

Test the selected model

Once choices are fixed, evaluate the selected model on an untouched test set reflecting intended use. Inspect errors by class and photo conditions, including lighting and background. Training vs Testing Data explains these separate roles.

Report the amount and coverage of test data alongside results. A small sample leaves substantial uncertainty. If test errors inspire further changes, that test set has become part of development; use fresh held-out evidence for the revised model’s final assessment.

Training fits parameters, validation guides choices, and a separate test set evaluates the selected model.
Keep related photos together and the test set outside model selection. View the diagram at full size.

When transfer learning helps and when it fails

Transfer learning is worth trying when source representations suit the target problem, particularly when target labels are limited. A general-photo backbone may already supply useful visual features, leaving less to learn from the shop’s examples. This can reduce data or training requirements, but the benefit depends on the source model and target task.

Domain mismatch can weaken that benefit. A model adapted only on clean studio photos may struggle with outdoor phone photos. Even with matching labels, backgrounds, lighting and camera quality can change. Representative target data matters whichever approach you choose.

Studio and outdoor bag photos differ in background and lighting while sharing the same class.
Changed photo conditions can weaken transfer; evaluate representative target images. View the diagram at full size.

Fine-tuning also creates room to overfit: the model can fit training examples without improving performance on new ones. Training improvement alongside worsening validation results is a warning sign, explored in Overfitting vs Underfitting.

Negative transfer means that using source knowledge harms performance relative to a suitable no-transfer comparison under comparable conditions. Wang and colleagues' research on negative transfer emphasizes this comparison. A weak result alone does not establish it; poor labels or an unsuitable training procedure might affect both methods.

Cost also needs measurement. A frozen backbone avoids its parameter updates, but still performs forward computation unless features are precomputed and reused. Fine-tuning more layers can increase memory and training cost. Neither transfer approach automatically makes inference faster. Compare quality, development effort and running cost under stated conditions.

How this relates to language models

Language models also reuse pretrained representations. Further training can adapt a model to a domain or task, although the data, objectives and model design differ from our image classifier. What Is GPT explains the pretraining foundation.

Ordinary text prompting supplies instructions or examples at inference time. Retrieval-augmented generation, or RAG, supplies retrieved information as context. Neither operation by itself updates model weights, so neither is fine-tuning. A RAG system can separately contain trained or fine-tuned components. Deciding whether to train requires its own data and evaluation plan.

Frequently asked questions

Is transfer learning the same as fine-tuning?

No. Transfer learning is the broader reuse strategy. Fine-tuning adapts pretrained parameters through further training. Training a new head while keeping pretrained features fixed is another transfer approach.

Is anything trained when the feature extractor is frozen?

Yes. In our example, the new head learns from target labels. Freezing the backbone prevents its parameter updates; it does not remove the head’s training process.

Does fine-tuning always update every layer?

No. You can update selected layers or the whole pretrained backbone. The choice should follow validation results, available data and practical constraints.

Must the new task use the original classes?

No. A replacement head can learn different target classes. Our general-object source model becomes a shoes, bags and hats classifier through a new head and target training.

How much target data is enough?

There is no universal image count. Difficulty, label quality, class coverage and source similarity matter. Check performance on representative held-out examples and gather data that addresses observed gaps.

Can transfer learning perform worse than training from scratch?

Yes. Mismatched source features or an unsuitable adaptation procedure can hurt. Use a properly trained no-transfer baseline to investigate whether transfer helps under your conditions.

Why are validation and test data kept separate?

Validation guides model choices. The test set assesses the selected procedure after those choices are fixed. Repeatedly choosing models using test results weakens that independent assessment.

Are ordinary prompting and RAG forms of fine-tuning?

No. They change the context supplied during inference without themselves updating learned weights. Fine-tuning requires a training process; a larger system may use both.

Does transfer learning work beyond computer vision?

Yes. It is also used in language, speech and other areas. What can be reused depends on the model, data and relationship between source and target problems.

Key takeaway and where to learn next

Transfer learning starts with learned knowledge, then asks which parts to reuse and which to train. Compare alternatives, validate your choices and test the selected model on relevant unseen data. Continue with Model Evaluation Metrics Explained to choose checks that match your task.

Scroll to Top