Few things in machine learning are as disappointing as a model that scores 99% in your notebook and fails in the real world. Very often the cause is data leakage: somewhere, information that the model should not have had slipped into training or evaluation.
A definition you can apply
Ask one question about every feature and every processing step:
Would this information exist, in exactly this form, at the moment the model makes a real prediction?
If the answer is no, you have leakage.
Common types of leakage
1. Preprocessing on the full dataset
Fitting a scaler, imputer, vocabulary, or feature selector before splitting lets statistics from the validation and test data shape training.
# Leaky: the scaler has seen the test data
X_scaled = StandardScaler().fit_transform(X)
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y)
# Correct: split first, and let a pipeline fit preprocessing on training data only
X_train, X_test, y_train, y_test = train_test_split(X, y, stratify=y, random_state=42)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X_train, y_train)With scaling the effect is often small. With feature selection or text vocabularies it can be large, because the selected features were chosen partly for how well they separate the test labels.
2. Duplicates and near-duplicates across splits
Security datasets are full of near-copies: the same phishing template sent to thousands of people with a different name, or the same malware sample repackaged. If copies land in both training and test sets, the model is tested on examples it has effectively already seen.
Fix: deduplicate before splitting, and for templated data, split by campaign or family rather than by individual row.
3. Group leakage
Many rows can come from one user, one device, one sender domain or one network host. A random split puts the same group on both sides, and the model learns to recognise the group instead of the pattern. Use group-aware splitting such as GroupKFold; see cross-validation.
4. Temporal leakage
If the data changes over time, a random split lets the model train on the future and test on the past. Attackers change tactics; a phishing detector trained on 2025 and 2026 emails and tested on 2024 emails will look better than it will perform next month. Split by time.
5. Target leakage (features that encode the answer)
A feature can quietly contain the label:
- A "quarantine_folder" field that is only set after an email was already judged malicious.
- An "incident_ticket_id" that exists only for confirmed incidents.
- A field produced by the same rule or analyst whose decision is the label.
These features are extremely predictive in your dataset and useless in reality, because they do not exist at prediction time.
6. Leakage through repeated tuning
Checking the test set, adjusting, and checking again gradually fits your choices to the test set. See train, validation and test sets.
Warning signs
- Performance that seems too good for a hard problem.
- One feature with overwhelming importance. Investigate it first.
- Validation or test scores higher than training scores.
- A big drop when you test on data collected later or from a different source.
Leakage in RAG and LLM evaluation
The same ideas apply to LLM systems:
- Test questions written while looking at the retrieved chunks tend to reuse the chunks' exact wording, which makes retrieval look easier than real user questions.
- Benchmark contamination: public test questions may have appeared in a model's training data, so high scores can reflect memory rather than ability.
- Tuning prompts on the evaluation set is the prompt-engineering version of tuning on the test set.
Keep a held-out set of real questions, as described in evaluating RAG systems.
Checklist
- Deduplicate, then split.
- Split by time, group, campaign or family when rows are not independent.
- Put every learned preprocessing step inside a pipeline.
- For each feature, confirm it exists at prediction time.
- Touch the test set once.
Summary
Data leakage makes evaluation scores measure the wrong thing. It is prevented less by clever code than by discipline: split correctly, fit preprocessing on training data only, and question any feature or score that looks too good.