Packet To Sniff

Cross-validation explained: k-fold, stratified and grouped

Cross-validation gives a more reliable performance estimate than one split by rotating the validation fold. Learn k-fold, stratified, group and time-series variants.

By Packet To SniffPublished 3 min read

Helpful background: Train, validation and test sets explained, Overfitting and underfitting in machine learning.

On this page

A single validation split gives you one number, and that number depends on which rows happened to land in it. With a small dataset, a lucky split can make a weak model look good. Cross-validation fixes this by validating on every part of the data in turn.

k-fold cross-validation, step by step

With k = 5:

  1. Shuffle the development data (everything except the test set) and split it into 5 equal folds.
  2. Train on folds 2 to 5, validate on fold 1.
  3. Train on folds 1, 3, 4, 5, validate on fold 2.
  4. Continue until each fold has been the validation fold once.
  5. Report the mean and the standard deviation of the 5 scores.
Output
Round 1: [VAL] [trn] [trn] [trn] [trn]
Round 2: [trn] [VAL] [trn] [trn] [trn]
Round 3: [trn] [trn] [VAL] [trn] [trn]
Round 4: [trn] [trn] [trn] [VAL] [trn]
Round 5: [trn] [trn] [trn] [trn] [VAL]

Every row is used for validation exactly once and for training k - 1 times.

In scikit-learn

Python
from sklearn.model_selection import StratifiedKFold, cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
 
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
 
scores = cross_val_score(model, X_dev, y_dev, cv=cv, scoring="f1")
print(f"F1: {scores.mean():.3f} ± {scores.std():.3f}")

Two details matter here:

  • The scaler is inside the pipeline, so in each round it is fitted only on that round's training folds. Scaling before cross-validation would leak validation information into training. See data leakage.
  • scoring="f1" reports the metric you care about. For imbalanced problems, accuracy is usually the wrong choice; see precision, recall and F1.

Choosing the right splitter

Your data Use Why
Classification with imbalanced classes StratifiedKFold Keeps class ratios equal in every fold
Several rows per user, device or thread GroupKFold or StratifiedGroupKFold Keeps each group in one fold only
Ordered in time TimeSeriesSplit Always validates on data later than training
Very small dataset RepeatedStratifiedKFold Repeats with different shuffles to steady the estimate

Using plain k-fold on grouped or time-ordered data produces scores that look better than reality.

Read the spread, not just the mean

Two models:

  • Model A: F1 = 0.84 ± 0.01
  • Model B: F1 = 0.85 ± 0.06

Model B's mean is higher, but its folds vary a lot. The difference between them is well inside B's spread, so you do not have good evidence that B is better. A large standard deviation is also a hint to look at the individual folds: one bad fold may reveal a subgroup where the model fails.

Tuning hyperparameters with cross-validation

GridSearchCV and RandomizedSearchCV run cross-validation for every setting you want to compare and pick the best:

Python
from sklearn.model_selection import GridSearchCV
 
search = GridSearchCV(
    model,
    param_grid={"logisticregression__C": [0.01, 0.1, 1, 10]},
    cv=cv, scoring="f1",
)
search.fit(X_dev, y_dev)
print(search.best_params_, round(search.best_score_, 3))

When not to use it

Cross-validation multiplies training time by k. For large deep-learning models that take hours or days to train, a single, large, well-constructed validation set is usually the practical choice.

Summary

  • k-fold cross-validation validates on every fold once and reports a mean and spread.
  • Pick the splitter that matches your data: stratified, grouped or time-based.
  • Keep preprocessing inside a pipeline, and still keep a separate test set for the final number.

Frequently asked questions

How many folds should I use?

Five or ten folds are common choices. More folds use more data for training in each round but take longer to run and can give more variable fold scores. With very small datasets, repeated k-fold cross-validation can steady the estimate further.

Does cross-validation replace the test set?

No. Cross-validation is for development decisions such as choosing models and hyperparameters. You still keep a separate test set untouched until the end to report the final result.

Continue learning

Tags

  • How to evaluate a RAG system: retrieval and answer quality

    Measure a RAG system in two layers: did retrieval find the right passages, and is the answer faithful to them? Learn recall@k, MRR, faithfulness and a practical test set.

    Retrieval-augmented generationIntermediate3 min
  • Why long context windows still miss things

    A million-token context window does not mean a model uses every token well. Learn about the lost-in-the-middle effect, attention cost, and how to test it yourself.

    LLM foundationsIntermediate3 min