Articles on AI, machine learning and security
Every article opens with a short answer, then explains the mechanism with examples you can run.
Cross-validation explained: k-fold, stratified and grouped
Cross-validation gives a more reliable performance estimate than one split by rotating the validation fold. Learn k-fold, stratified, group and time-series variants.
Machine learning methodsIntermediate3 minData leakage in machine learning: how good scores lie
Data leakage lets outside information slip into training, inflating scores that collapse in production. Learn the common types and how to prevent them.
Machine learning methodsIntermediate3 minHow to chunk documents for RAG
Chunking decides what your RAG system can retrieve. Compare fixed-size, recursive, heading-aware and semantic chunking, and learn how to pick size and overlap.
Retrieval-augmented generationIntermediate3 minHow to evaluate a RAG system: retrieval and answer quality
Measure a RAG system in two layers: did retrieval find the right passages, and is the answer faithful to them? Learn recall@k, MRR, faithfulness and a practical test set.
Retrieval-augmented generationIntermediate3 minHow vector search works: similarity, top-k and hybrid search
Vector search finds text by meaning instead of exact words. Learn cosine similarity, top-k retrieval, approximate nearest neighbour indexes, BM25 and hybrid ranking.
Retrieval-augmented generationIntermediate3 minHTTP security headers explained: CSP, HSTS and more
Security headers tell browsers how to protect your users. Learn what CSP, HSTS, X-Content-Type-Options, Referrer-Policy, Permissions-Policy and frame-ancestors do.
Web securityBeginner3 minOverfitting and underfitting in machine learning
Overfitting means a model memorises its training data; underfitting means it misses the pattern. Learn to spot both from your scores and how to fix each.
Machine learning methodsBeginner3 minPrecision, recall and F1 score explained with an example
Accuracy can be misleading on imbalanced data. Learn the confusion matrix, precision, recall, F1 and thresholds through a worked phishing-detection example.
Machine learning methodsBeginner3 minPrompt injection in RAG systems and how to reduce it
Prompt injection makes a language model follow attacker-written text. Learn direct and indirect injection in RAG, why filters fail, and layered defences.
AI securityIntermediate3 minTCP vs UDP: differences, handshakes and when each is used
TCP gives reliable, ordered delivery after a handshake; UDP sends independent datagrams. Compare headers, use cases, ports and how both look in Wireshark.
NetworkingBeginner3 minTrain, validation and test sets explained
Why machine learning data is split three ways, what each split is for, how to split correctly with scikit-learn, and the mistakes that make a test score meaningless.
Machine learning methodsBeginner3 minWhat are embeddings? Meaning as numbers, explained
Embeddings turn text into lists of numbers so that similar meanings land close together. Learn how they work, how similarity is measured, and where they fail.
LLM foundationsBeginner3 minWhat is a context window in an LLM?
The context window is the maximum number of tokens a language model can consider at once. Learn what fills it, what happens when it overflows, and how to budget it.
LLM foundationsBeginner3 minWhat is a token in an LLM?
A token is the unit of text a language model reads and writes. Learn how tokenizers split text, why token counts differ by language, and why it matters for cost.
LLM foundationsBeginner4 minWhat is retrieval-augmented generation (RAG)?
RAG makes a language model answer from your own documents by retrieving relevant passages first. Learn the pipeline step by step, and when RAG is the right choice.
Retrieval-augmented generationBeginner3 minWhy long context windows still miss things
A million-token context window does not mean a model uses every token well. Learn about the lost-in-the-middle effect, attention cost, and how to test it yourself.
LLM foundationsIntermediate3 min