Packet To Sniff

Data leakage in machine learning: how good scores lie

Data leakage lets outside information slip into training, inflating scores that collapse in production. Learn the common types and how to prevent them.

By Packet To SniffPublished 3 min read

Helpful background: Train, validation and test sets explained, Cross-validation explained: k-fold, stratified and grouped.

On this page

Few things in machine learning are as disappointing as a model that scores 99% in your notebook and fails in the real world. Very often the cause is data leakage: somewhere, information that the model should not have had slipped into training or evaluation.

A definition you can apply

Ask one question about every feature and every processing step:

Would this information exist, in exactly this form, at the moment the model makes a real prediction?

If the answer is no, you have leakage.

Common types of leakage

1. Preprocessing on the full dataset

Fitting a scaler, imputer, vocabulary, or feature selector before splitting lets statistics from the validation and test data shape training.

Python
# Leaky: the scaler has seen the test data
X_scaled = StandardScaler().fit_transform(X)
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y)
 
# Correct: split first, and let a pipeline fit preprocessing on training data only
X_train, X_test, y_train, y_test = train_test_split(X, y, stratify=y, random_state=42)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X_train, y_train)

With scaling the effect is often small. With feature selection or text vocabularies it can be large, because the selected features were chosen partly for how well they separate the test labels.

2. Duplicates and near-duplicates across splits

Security datasets are full of near-copies: the same phishing template sent to thousands of people with a different name, or the same malware sample repackaged. If copies land in both training and test sets, the model is tested on examples it has effectively already seen.

Fix: deduplicate before splitting, and for templated data, split by campaign or family rather than by individual row.

3. Group leakage

Many rows can come from one user, one device, one sender domain or one network host. A random split puts the same group on both sides, and the model learns to recognise the group instead of the pattern. Use group-aware splitting such as GroupKFold; see cross-validation.

4. Temporal leakage

If the data changes over time, a random split lets the model train on the future and test on the past. Attackers change tactics; a phishing detector trained on 2025 and 2026 emails and tested on 2024 emails will look better than it will perform next month. Split by time.

5. Target leakage (features that encode the answer)

A feature can quietly contain the label:

  • A "quarantine_folder" field that is only set after an email was already judged malicious.
  • An "incident_ticket_id" that exists only for confirmed incidents.
  • A field produced by the same rule or analyst whose decision is the label.

These features are extremely predictive in your dataset and useless in reality, because they do not exist at prediction time.

6. Leakage through repeated tuning

Checking the test set, adjusting, and checking again gradually fits your choices to the test set. See train, validation and test sets.

Warning signs

  • Performance that seems too good for a hard problem.
  • One feature with overwhelming importance. Investigate it first.
  • Validation or test scores higher than training scores.
  • A big drop when you test on data collected later or from a different source.

Leakage in RAG and LLM evaluation

The same ideas apply to LLM systems:

  • Test questions written while looking at the retrieved chunks tend to reuse the chunks' exact wording, which makes retrieval look easier than real user questions.
  • Benchmark contamination: public test questions may have appeared in a model's training data, so high scores can reflect memory rather than ability.
  • Tuning prompts on the evaluation set is the prompt-engineering version of tuning on the test set.

Keep a held-out set of real questions, as described in evaluating RAG systems.

Checklist

  1. Deduplicate, then split.
  2. Split by time, group, campaign or family when rows are not independent.
  3. Put every learned preprocessing step inside a pipeline.
  4. For each feature, confirm it exists at prediction time.
  5. Touch the test set once.

Summary

Data leakage makes evaluation scores measure the wrong thing. It is prevented less by clever code than by discipline: split correctly, fit preprocessing on training data only, and question any feature or score that looks too good.

Frequently asked questions

How can I tell if my model has data leakage?

Warning signs include scores that seem too good for the problem, a single feature with extreme importance, validation scores higher than training scores, and a large drop in performance on genuinely new data. Check for duplicates across splits and ask of every feature whether it would really exist at prediction time.

Is data leakage a security problem?

The term also has a security meaning: sensitive data leaking out of a system. In machine learning evaluation it means information leaking into training. Both matter for AI systems, but they are different problems with different fixes.

Continue learning

Tags

  • Overfitting and underfitting in machine learning

    Overfitting means a model memorises its training data; underfitting means it misses the pattern. Learn to spot both from your scores and how to fix each.

    Machine learning methodsBeginner3 min
  • Why long context windows still miss things

    A million-token context window does not mean a model uses every token well. Learn about the lost-in-the-middle effect, attention cost, and how to test it yourself.

    LLM foundationsIntermediate3 min