A single validation split gives you one number, and that number depends on which rows happened to land in it. With a small dataset, a lucky split can make a weak model look good. Cross-validation fixes this by validating on every part of the data in turn.
k-fold cross-validation, step by step
With k = 5:
- Shuffle the development data (everything except the test set) and split it into 5 equal folds.
- Train on folds 2 to 5, validate on fold 1.
- Train on folds 1, 3, 4, 5, validate on fold 2.
- Continue until each fold has been the validation fold once.
- Report the mean and the standard deviation of the 5 scores.
Round 1: [VAL] [trn] [trn] [trn] [trn]
Round 2: [trn] [VAL] [trn] [trn] [trn]
Round 3: [trn] [trn] [VAL] [trn] [trn]
Round 4: [trn] [trn] [trn] [VAL] [trn]
Round 5: [trn] [trn] [trn] [trn] [VAL]Every row is used for validation exactly once and for training k - 1 times.
In scikit-learn
from sklearn.model_selection import StratifiedKFold, cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X_dev, y_dev, cv=cv, scoring="f1")
print(f"F1: {scores.mean():.3f} ± {scores.std():.3f}")Two details matter here:
- The scaler is inside the pipeline, so in each round it is fitted only on that round's training folds. Scaling before cross-validation would leak validation information into training. See data leakage.
scoring="f1"reports the metric you care about. For imbalanced problems, accuracy is usually the wrong choice; see precision, recall and F1.
Choosing the right splitter
| Your data | Use | Why |
|---|---|---|
| Classification with imbalanced classes | StratifiedKFold |
Keeps class ratios equal in every fold |
| Several rows per user, device or thread | GroupKFold or StratifiedGroupKFold |
Keeps each group in one fold only |
| Ordered in time | TimeSeriesSplit |
Always validates on data later than training |
| Very small dataset | RepeatedStratifiedKFold |
Repeats with different shuffles to steady the estimate |
Using plain k-fold on grouped or time-ordered data produces scores that look better than reality.
Read the spread, not just the mean
Two models:
- Model A: F1 = 0.84 ± 0.01
- Model B: F1 = 0.85 ± 0.06
Model B's mean is higher, but its folds vary a lot. The difference between them is well inside B's spread, so you do not have good evidence that B is better. A large standard deviation is also a hint to look at the individual folds: one bad fold may reveal a subgroup where the model fails.
Tuning hyperparameters with cross-validation
GridSearchCV and RandomizedSearchCV run cross-validation for every setting you want to compare and pick the best:
from sklearn.model_selection import GridSearchCV
search = GridSearchCV(
model,
param_grid={"logisticregression__C": [0.01, 0.1, 1, 10]},
cv=cv, scoring="f1",
)
search.fit(X_dev, y_dev)
print(search.best_params_, round(search.best_score_, 3))When not to use it
Cross-validation multiplies training time by k. For large deep-learning models that take hours or days to train, a single, large, well-constructed validation set is usually the practical choice.
Summary
- k-fold cross-validation validates on every fold once and reports a mean and spread.
- Pick the splitter that matches your data: stratified, grouped or time-based.
- Keep preprocessing inside a pipeline, and still keep a separate test set for the final number.