Packet To Sniff

Precision, recall and F1 score explained with an example

Accuracy can be misleading on imbalanced data. Learn the confusion matrix, precision, recall, F1 and thresholds through a worked phishing-detection example.

By Packet To SniffPublished 3 min read

Helpful background: Train, validation and test sets explained.

On this page

Suppose you build a model to flag phishing emails and it reports 95.4% accuracy. Sounds great. But if 90% of emails are legitimate anyway, a model that never flags anything scores 90%. Accuracy hides the thing you care about. Precision and recall show it.

The confusion matrix

Every prediction from a binary classifier falls into one of four boxes:

Predicted phishing Predicted legitimate
Actually phishing True positive (TP) False negative (FN)
Actually legitimate False positive (FP) True negative (TN)
  • False positive: a legitimate email flagged as phishing (a false alarm).
  • False negative: a phishing email that got through (a miss).

A worked example

Take 1,000 emails, of which 100 are phishing. The model flags 90 emails, and 72 of those really are phishing.

Predicted phishing Predicted legitimate Total
Actually phishing TP = 72 FN = 28 100
Actually legitimate FP = 18 TN = 882 900

Now compute each metric.

Accuracy: correct predictions over all predictions.

Output
(TP + TN) / total = (72 + 882) / 1000 = 0.954

Precision: of everything flagged, how much was really phishing?

Output
TP / (TP + FP) = 72 / (72 + 18) = 72 / 90 = 0.80

Recall: of all real phishing, how much did we catch?

Output
TP / (TP + FN) = 72 / (72 + 28) = 72 / 100 = 0.72

F1 score: the harmonic mean of precision and recall.

Output
2 × P × R / (P + R) = 2 × 0.80 × 0.72 / (0.80 + 0.72) = 1.152 / 1.52 ≈ 0.758

So the "95.4% accurate" model actually lets 28% of phishing emails through. That is the number a security team needs to see.

Computing it in scikit-learn

Python
from sklearn.metrics import confusion_matrix, classification_report
 
# y_true and y_pred: 1 = phishing, 0 = legitimate
print(confusion_matrix(y_true, y_pred, labels=[1, 0]))
print(classification_report(y_true, y_pred, labels=[1, 0],
                            target_names=["phishing", "legitimate"], digits=3))

classification_report prints precision, recall and F1 for each class, plus macro and weighted averages.

Thresholds: trading precision for recall

Most classifiers output a probability, and a threshold (often 0.5 by default) turns it into a label. Moving the threshold trades one metric for the other:

  • Lower threshold (flag more): recall goes up, precision usually goes down.
  • Higher threshold (flag less): precision goes up, recall usually goes down.
Python
from sklearn.metrics import precision_recall_curve
 
probs = model.predict_proba(X_val)[:, 1]
precision, recall, thresholds = precision_recall_curve(y_val, probs)
# Pick the threshold that meets your requirement, e.g. recall >= 0.95,
# using the VALIDATION set, then report final numbers on the TEST set.

The right threshold is a business and security decision. What does a missed phishing email cost compared with a legitimate email sitting in quarantine? Choose the threshold on validation data, never on the test set.

Which metric should you report?

Situation Focus on
Misses are dangerous (phishing, fraud, intrusion) Recall, at an acceptable precision
False alarms are expensive or annoying Precision
You need one number to compare models F1, or area under the precision-recall curve
Classes are balanced and errors cost the same Accuracy is acceptable

Always report the confusion matrix too. It lets readers compute any metric they need.

Summary

  • Accuracy is misleading when one class is rare.
  • Precision asks "how many flags were right?"; recall asks "how many real cases did we catch?"
  • F1 balances the two; the threshold lets you choose where on the trade-off to sit.
  • Pick thresholds on validation data and report test results once.

Frequently asked questions

When should I prefer precision over recall?

Prefer precision when false alarms are expensive, for example blocking legitimate emails that a business depends on. Prefer recall when misses are expensive, for example letting a phishing email reach an employee. Many real systems choose a threshold that balances both for their costs.

Why is accuracy misleading for imbalanced data?

If only 10 percent of emails are phishing, a model that labels every email as legitimate is 90 percent accurate while catching zero phishing emails. Accuracy is dominated by the majority class.

What is a good F1 score?

There is no universal threshold. It depends on the task, the class balance and the cost of errors. Compare against a simple baseline on the same data and decide based on what false positives and false negatives cost in practice.

Continue learning

Tags

  • Overfitting and underfitting in machine learning

    Overfitting means a model memorises its training data; underfitting means it misses the pattern. Learn to spot both from your scores and how to fix each.

    Machine learning methodsBeginner3 min
  • How to evaluate a RAG system: retrieval and answer quality

    Measure a RAG system in two layers: did retrieval find the right passages, and is the answer faithful to them? Learn recall@k, MRR, faithfulness and a practical test set.

    Retrieval-augmented generationIntermediate3 min
  • Why long context windows still miss things

    A million-token context window does not mean a model uses every token well. Learn about the lost-in-the-middle effect, attention cost, and how to test it yourself.

    LLM foundationsIntermediate3 min