A model that scores 99% on the data it trained on has proved almost nothing. It may have simply memorised that data. What we actually care about is how well it does on new data, and the only honest way to estimate that is to hide some data from the model and from ourselves.
The three splits
| Split | Used for | Touched how often |
|---|---|---|
| Training | Fitting the model's parameters | Constantly |
| Validation | Comparing models, tuning hyperparameters, early stopping | Many times during development |
| Test | One final, unbiased estimate | Once, at the end |
A useful way to think about it:
- The training set is the textbook.
- The validation set is the practice exam you take repeatedly while studying.
- The test set is the real exam. If you have seen it before, your score is not a real score.
Why validation and test must be separate
Every time you look at a score and change something (a model type, a learning rate, a feature), you are making a decision based on that data. Make enough decisions and your choices start fitting the quirks of the validation set, even though no model was ever trained on it directly.
That is fine, as long as a separate test set stays untouched. The test score then answers the question: how will this final, chosen model do on data none of my decisions were based on?
Splitting in scikit-learn
train_test_split only makes two parts, so call it twice. Use stratify for classification and a fixed random_state so the split is reproducible.
from sklearn.model_selection import train_test_split
# X: features, y: labels (e.g. 1 = phishing, 0 = legitimate)
X_trainval, X_test, y_trainval, y_test = train_test_split(
X, y, test_size=0.15, stratify=y, random_state=42
)
X_train, X_val, y_train, y_val = train_test_split(
X_trainval, y_trainval,
test_size=0.15 / 0.85, # 15% of the original data
stratify=y_trainval, random_state=42,
)
print(len(X_train), len(X_val), len(X_test))With small datasets, a single validation split is noisy. Cross-validation gives a steadier estimate by rotating which part is used for validation.
Not every dataset should be shuffled
Random splits assume each example is independent. Often they are not:
- Time series and anything that changes over time. Train on the past, validate and test on the future. A random split lets the model "see the future", which inflates scores. Phishing campaigns, prices and network traffic all drift over time.
- Groups. If one user, patient, device or email thread appears many times, put all its rows in the same split. Otherwise the model can recognise the group instead of learning the pattern. scikit-learn's
GroupShuffleSplitandGroupKFoldhandle this. - Near-duplicates. Scraped datasets often contain copies with tiny differences. Remove duplicates before splitting, or the same example can appear in train and test.
These are all forms of data leakage, the most common reason a model's reported score does not survive the real world.
Preprocessing belongs after the split
Anything learned from data (scaling means, vocabulary, feature selection, imputation values) must be learned from the training data only, then applied to validation and test. Fitting a scaler on the full dataset before splitting leaks information about the test set into training. Pipelines make this automatic:
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_train, y_train) # scaler learns from training data only
print(model.score(X_val, y_val)) # validation accuracySummary
- Train on the training set, make decisions with the validation set, and report the test set once.
- Stratify for imbalanced classes; split by time or group when rows are not independent.
- Fit all preprocessing on training data only, ideally inside a pipeline.
Next, see what the gap between training and validation scores tells you in Overfitting and underfitting.