Machine learning methods
Data splits, overfitting, cross-validation and evaluation metrics: the habits that decide whether a model's score means anything.
Part of Artificial intelligence. 5 articles.
Cross-validation explained: k-fold, stratified and grouped
Cross-validation gives a more reliable performance estimate than one split by rotating the validation fold. Learn k-fold, stratified, group and time-series variants.
Intermediate3 minData leakage in machine learning: how good scores lie
Data leakage lets outside information slip into training, inflating scores that collapse in production. Learn the common types and how to prevent them.
Intermediate3 minOverfitting and underfitting in machine learning
Overfitting means a model memorises its training data; underfitting means it misses the pattern. Learn to spot both from your scores and how to fix each.
Beginner3 minPrecision, recall and F1 score explained with an example
Accuracy can be misleading on imbalanced data. Learn the confusion matrix, precision, recall, F1 and thresholds through a worked phishing-detection example.
Beginner3 minTrain, validation and test sets explained
Why machine learning data is split three ways, what each split is for, how to split correctly with scikit-learn, and the mistakes that make a test score meaningless.
Beginner3 min