Cross-Validation Explained: A Beginner-Friendly Guide

Cross-validation evaluates a model across several training and validation splits. It helps you compare model choices and see how much the results change when different examples are held out. It is especially useful when a single validation split would make the decision depend too heavily on which examples happened to land in that subset.

The key is to test the whole learning process fairly: choose splits that fit the data, learn preprocessing inside each training fold, and keep the final test set separate from development decisions.

How k-fold cross-validation works

In k-fold cross-validation, you divide development data into k groups called folds. Train a fresh model on k minus one folds, evaluate it on the remaining fold, and repeat until each fold has served as validation data once. The held-out fold changes each round; the model does not carry learned information from an earlier round into the next.

For an illustrative five-fold example, suppose you have 1,000 development examples after reserving a separate final test set. Divide those examples into five folds of 200. Each round trains on 800 examples and validates on 200:

An illustrative five-fold rotation
Round Training folds Validation fold
1 B, C, D, E A
2 A, C, D, E B
3 A, B, D, E C
4 A, B, C, E D
5 A, B, C, D E

Every development example is used for validation once and for training in the other four rounds. The separate final test set is absent from all five rounds. See the scikit-learn cross-validation guide for the underlying workflow.

Five cross-validation rounds rotate the validation role across folds A to E while the other four folds train the model; the final test set stays separate.
Each development fold is used for validation once. The independent final test set does not take part in these rounds. View the diagram at full size.

Read the average and the variation

Suppose five illustrative accuracy scores are 78%, 80%, 83%, 79%, and 80%. Their mean is 80%, and the scores range from 78% to 83%. Reporting both tells you more than presenting the average alone. These numbers explain the calculation; they are not results from a real experiment.

Large differences between folds deserve investigation. Check which examples, groups, or periods are difficult and whether each fold has enough relevant cases. The spread across folds is a useful diagnostic, not automatically a confidence interval for future performance. Stratification can also make fold scores look more similar without eliminating uncertainty about rare classes.

Choose the metric for the problem before comparing candidates. Accuracy, recall, and a regression error measure answer different questions. The scikit-learn scoring guide explains metric selection; Model Evaluation Metrics Explained introduces the main choices.

Choose a validation strategy that fits the data

Ordinary k-fold

Ordinary k-fold is a starting point when examples can reasonably be treated as independent and the folds represent the prediction task. It does not preserve class ratios or keep related records together by itself. Decide whether shuffling is appropriate instead of applying it automatically.

Stratified k-fold

For classification, stratified k-fold keeps class proportions approximately similar across folds. This helps avoid poorly represented classes in individual splits. It cannot create missing examples, guarantee reliable evaluation of a tiny minority class, or fix group or time leakage. See StratifiedKFold.

Grouped cross-validation

If several records come from the same person, device, or organization, a row-level split may put closely related information on both sides. Grouped cross-validation keeps a group out of the corresponding training set when that group is being evaluated. For example, evaluating on unseen customers may require holding out entire customers. The choice should match what will be new at deployment. See GroupKFold.

Time-series validation

For forecasting, train on earlier observations and evaluate on later ones. Rolling or expanding windows can represent repeated predictions over time. A gap between training and evaluation may be needed when features or targets overlap in time. Use only information that would have been available at the prediction moment. See TimeSeriesSplit.

Keep preprocessing inside each training fold

Cross-validation does not repair information leakage. Scaling, imputation, feature selection, and other learned transformations should be fitted using only the training portion of each fold. Apply those fitted transformations to that fold’s validation data, then measure the predictions.

For each round: split the data, fit preprocessing on the training rows, fit the model on the transformed training rows, transform the validation rows with the training-fitted preprocessing, and score them. Repeat the entire process with a fresh pipeline for the next fold.

Computing feature-selection scores or imputation values from the full dataset before cross-validation can reveal information about validation examples. Pipelines help keep these steps together. See scikit-learn’s leakage guidance, Data Preprocessing Explained, and Feature Selection vs Feature Extraction.

Training data fits preprocessing and the model; validation data only receives the fitted transformations and predictions before scoring.
Fit the entire pipeline on each round’s training data, then apply it to that round’s held-out data. View the diagram at full size.

Use cross-validation for development and preserve a final test

Cross-validation can compare algorithms, preprocessing choices, feature sets, and hyperparameters. A hyperparameter is a setting chosen outside model fitting, such as the number of neighbors in a nearest-neighbors model. Apply the same appropriate splits and scoring rule to the candidates so comparisons are meaningful. See scikit-learn’s hyperparameter-tuning guide.

After selecting the approach, refit it on the development data and evaluate it on the untouched final test set when you need an independent final assessment. Do not keep adjusting the system in response to that test score. If the train, validation, and test roles are unfamiliar, start with Training vs Testing Data.

Cross-validation does not inherently prevent overfitting. Repeatedly choosing whichever experiment looks best can overfit the selection process itself. Cawley and Talbot’s study of model-selection bias explains this risk. For the broader learning problem, read Overfitting vs Underfitting.

Advanced note: nested cross-validation separates an inner loop for tuning from an outer loop for evaluating the tuning procedure. It requires more computation and careful splitting in both loops. See the nested cross-validation example.

Choose the number of folds thoughtfully

Five or ten folds are common starting points, but there is no universal best value. Consider dataset size, training cost, class counts, group structure, and the deployment question. More folds require more model fits and do not automatically make an unsuitable split reliable. With very rare classes or few independent groups, you may need fewer folds or more representative data.

A practical cross-validation checklist

  • Define what future examples, groups, or time periods the model must predict.
  • Choose a splitting strategy and useful evaluation metrics.
  • Reserve a final test set if an independent final assessment is required.
  • Fit preprocessing and feature selection inside each training fold.
  • Inspect fold scores and errors as well as the average.
  • Choose the model using development evidence, then evaluate the final approach without retuning on the test set.

Frequently asked questions

What is cross-validation in simple terms?

It evaluates a model using several training and validation splits so a development decision does not rely on just one split.

Is cross-validation the same as a final test set?

No. Cross-validation usually supports development and model selection. A final test set stays separate to assess the selected approach without using that result for further tuning.

Should time-series data use ordinary k-fold?

For forecasting future observations from past data, validation should preserve time order. Ordinary shuffled k-fold can put future information into training.

Does preprocessing happen before cross-validation?

Learned preprocessing should be fitted inside each training fold and then applied to its validation fold. Fitting it on the full dataset first can leak information.

Where to learn next

Continue with Model Evaluation Metrics Explained to understand the scores used in cross-validation. Revisit Training vs Testing Data if you need a refresher on the separate data roles.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top