Packet To Sniff

Train, validation and test sets explained

Why machine learning data is split three ways, what each split is for, how to split correctly with scikit-learn, and the mistakes that make a test score meaningless.

By Packet To SniffPublished 3 min read

On this page

A model that scores 99% on the data it trained on has proved almost nothing. It may have simply memorised that data. What we actually care about is how well it does on new data, and the only honest way to estimate that is to hide some data from the model and from ourselves.

The three splits

Split Used for Touched how often
Training Fitting the model's parameters Constantly
Validation Comparing models, tuning hyperparameters, early stopping Many times during development
Test One final, unbiased estimate Once, at the end

A useful way to think about it:

  • The training set is the textbook.
  • The validation set is the practice exam you take repeatedly while studying.
  • The test set is the real exam. If you have seen it before, your score is not a real score.

Why validation and test must be separate

Every time you look at a score and change something (a model type, a learning rate, a feature), you are making a decision based on that data. Make enough decisions and your choices start fitting the quirks of the validation set, even though no model was ever trained on it directly.

That is fine, as long as a separate test set stays untouched. The test score then answers the question: how will this final, chosen model do on data none of my decisions were based on?

Splitting in scikit-learn

train_test_split only makes two parts, so call it twice. Use stratify for classification and a fixed random_state so the split is reproducible.

Python
from sklearn.model_selection import train_test_split
 
# X: features, y: labels (e.g. 1 = phishing, 0 = legitimate)
X_trainval, X_test, y_trainval, y_test = train_test_split(
    X, y, test_size=0.15, stratify=y, random_state=42
)
X_train, X_val, y_train, y_val = train_test_split(
    X_trainval, y_trainval,
    test_size=0.15 / 0.85,        # 15% of the original data
    stratify=y_trainval, random_state=42,
)
print(len(X_train), len(X_val), len(X_test))

With small datasets, a single validation split is noisy. Cross-validation gives a steadier estimate by rotating which part is used for validation.

Not every dataset should be shuffled

Random splits assume each example is independent. Often they are not:

  1. Time series and anything that changes over time. Train on the past, validate and test on the future. A random split lets the model "see the future", which inflates scores. Phishing campaigns, prices and network traffic all drift over time.
  2. Groups. If one user, patient, device or email thread appears many times, put all its rows in the same split. Otherwise the model can recognise the group instead of learning the pattern. scikit-learn's GroupShuffleSplit and GroupKFold handle this.
  3. Near-duplicates. Scraped datasets often contain copies with tiny differences. Remove duplicates before splitting, or the same example can appear in train and test.

These are all forms of data leakage, the most common reason a model's reported score does not survive the real world.

Preprocessing belongs after the split

Anything learned from data (scaling means, vocabulary, feature selection, imputation values) must be learned from the training data only, then applied to validation and test. Fitting a scaler on the full dataset before splitting leaks information about the test set into training. Pipelines make this automatic:

Python
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
 
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_train, y_train)          # scaler learns from training data only
print(model.score(X_val, y_val))     # validation accuracy

Summary

  • Train on the training set, make decisions with the validation set, and report the test set once.
  • Stratify for imbalanced classes; split by time or group when rows are not independent.
  • Fit all preprocessing on training data only, ideally inside a pipeline.

Next, see what the gap between training and validation scores tells you in Overfitting and underfitting.

Frequently asked questions

What is a good train, validation and test split ratio?

A common starting point is 70/15/15 or 80/10/10. With very large datasets the validation and test sets can be a smaller percentage, because what matters is having enough examples to estimate performance reliably, not a specific ratio.

Why not just use the test set to tune the model?

Every decision you make by looking at test results fits the model a little to that test set. After many such decisions, the test score becomes optimistic and no longer estimates performance on genuinely new data.

When should I use a stratified split?

Use stratification for classification when classes are imbalanced, such as rare phishing emails among normal ones. It keeps the class proportions the same in every split, so a small split does not end up with very few or no examples of the rare class.

Continue learning

Tags