A model has two jobs that pull against each other: fit the training data well enough to capture the real pattern, and not so closely that it memorises accidents. Underfitting fails the first job. Overfitting fails the second.
The pattern in three pictures
Imagine fitting a line to points that follow a gentle curve with some noise.
Underfit Good fit Overfit
(too simple) (captures trend) (memorises noise)
. . . . . . . . .
.------- . . /‾‾‾\ . ./\/\/\/\.
. . . /. .\ . . .- The underfit model draws a straight line through a curve. Wrong everywhere, in the same way.
- The good model follows the trend and ignores the wiggles.
- The overfit model threads through every single point, including the noise. It is perfect on training data and erratic on anything new.
Read it from the scores
You need a training score and a validation score:
| Training score | Validation score | Diagnosis |
|---|---|---|
| Low | Low | Underfitting |
| High | High, close to training | Good fit |
| Very high | Much lower | Overfitting |
| Low | Higher than training | Check for a bug or leakage |
The last row is a warning sign. Validation should rarely beat training by much; when it does, look for data leakage, a mismatch between splits, or regularisation that is only active during training (such as dropout).
A concrete example with decision trees
Decision trees make this easy to see, because max_depth directly controls complexity.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = load_breast_cancer(return_X_y=True)
X_train, X_val, y_train, y_val = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=0
)
for depth in [1, 2, 4, 8, None]:
tree = DecisionTreeClassifier(max_depth=depth, random_state=0)
tree.fit(X_train, y_train)
print(depth, round(tree.score(X_train, y_train), 3), round(tree.score(X_val, y_val), 3))Typically you will see training accuracy climb to 1.0 as depth grows, while validation accuracy peaks at a moderate depth and then stops improving or drops. The unlimited tree (None) memorises the training set. Run it yourself in the lab Measure overfitting with learning curves.
Fixing underfitting
- Use a more expressive model or add features that carry real signal.
- Reduce regularisation.
- Train longer, if the training score is still improving.
- Check the data: labels that are mostly noise cannot be fitted by any model.
Fixing overfitting
- Get more, more varied data. The most reliable fix when possible.
- Simplify the model. Fewer layers, shallower trees, fewer features.
- Regularise. L2 or L1 penalties, dropout, limits on tree depth or leaf size.
- Stop early. Watch validation loss and stop when it starts rising.
- Augment data where it makes sense, for example small image transformations.
The bias-variance view
These two failures are often described as bias and variance:
- High bias: the model's assumptions are too rigid. It is consistently wrong in the same way. This is underfitting.
- High variance: the model is too sensitive to the particular training sample. Train it on a slightly different sample and you get a very different model. This is overfitting.
Increasing model complexity usually lowers bias and raises variance. The goal is the sweet spot for your amount of data.
The same idea in LLM systems
Overfitting is not only about model weights. If you tune a RAG prompt, chunk size or retrieval settings over and over against the same 20 questions, your system can overfit those questions. The fix is the same: keep a held-out set of questions you do not tune on. See evaluating RAG systems.
Summary
- Underfitting: poor on training and validation. Overfitting: great on training, worse on validation.
- Compare training and validation scores to diagnose; fix underfitting with capacity and overfitting with data, simplicity or regularisation.
- Any system tuned repeatedly on the same examples can overfit them.