Suppose you build a model to flag phishing emails and it reports 95.4% accuracy. Sounds great. But if 90% of emails are legitimate anyway, a model that never flags anything scores 90%. Accuracy hides the thing you care about. Precision and recall show it.
The confusion matrix
Every prediction from a binary classifier falls into one of four boxes:
| Predicted phishing | Predicted legitimate | |
|---|---|---|
| Actually phishing | True positive (TP) | False negative (FN) |
| Actually legitimate | False positive (FP) | True negative (TN) |
- False positive: a legitimate email flagged as phishing (a false alarm).
- False negative: a phishing email that got through (a miss).
A worked example
Take 1,000 emails, of which 100 are phishing. The model flags 90 emails, and 72 of those really are phishing.
| Predicted phishing | Predicted legitimate | Total | |
|---|---|---|---|
| Actually phishing | TP = 72 | FN = 28 | 100 |
| Actually legitimate | FP = 18 | TN = 882 | 900 |
Now compute each metric.
Accuracy: correct predictions over all predictions.
(TP + TN) / total = (72 + 882) / 1000 = 0.954Precision: of everything flagged, how much was really phishing?
TP / (TP + FP) = 72 / (72 + 18) = 72 / 90 = 0.80Recall: of all real phishing, how much did we catch?
TP / (TP + FN) = 72 / (72 + 28) = 72 / 100 = 0.72F1 score: the harmonic mean of precision and recall.
2 × P × R / (P + R) = 2 × 0.80 × 0.72 / (0.80 + 0.72) = 1.152 / 1.52 ≈ 0.758So the "95.4% accurate" model actually lets 28% of phishing emails through. That is the number a security team needs to see.
Computing it in scikit-learn
from sklearn.metrics import confusion_matrix, classification_report
# y_true and y_pred: 1 = phishing, 0 = legitimate
print(confusion_matrix(y_true, y_pred, labels=[1, 0]))
print(classification_report(y_true, y_pred, labels=[1, 0],
target_names=["phishing", "legitimate"], digits=3))classification_report prints precision, recall and F1 for each class, plus macro and weighted averages.
Thresholds: trading precision for recall
Most classifiers output a probability, and a threshold (often 0.5 by default) turns it into a label. Moving the threshold trades one metric for the other:
- Lower threshold (flag more): recall goes up, precision usually goes down.
- Higher threshold (flag less): precision goes up, recall usually goes down.
from sklearn.metrics import precision_recall_curve
probs = model.predict_proba(X_val)[:, 1]
precision, recall, thresholds = precision_recall_curve(y_val, probs)
# Pick the threshold that meets your requirement, e.g. recall >= 0.95,
# using the VALIDATION set, then report final numbers on the TEST set.The right threshold is a business and security decision. What does a missed phishing email cost compared with a legitimate email sitting in quarantine? Choose the threshold on validation data, never on the test set.
Which metric should you report?
| Situation | Focus on |
|---|---|
| Misses are dangerous (phishing, fraud, intrusion) | Recall, at an acceptable precision |
| False alarms are expensive or annoying | Precision |
| You need one number to compare models | F1, or area under the precision-recall curve |
| Classes are balanced and errors cost the same | Accuracy is acceptable |
Always report the confusion matrix too. It lets readers compute any metric they need.
Summary
- Accuracy is misleading when one class is rare.
- Precision asks "how many flags were right?"; recall asks "how many real cases did we catch?"
- F1 balances the two; the threshold lets you choose where on the trade-off to sit.
- Pick thresholds on validation data and report test results once.