Packet To Sniff

Articles on AI, machine learning and security

Every article opens with a short answer, then explains the mechanism with examples you can run.

  • Cross-validation explained: k-fold, stratified and grouped

    Cross-validation gives a more reliable performance estimate than one split by rotating the validation fold. Learn k-fold, stratified, group and time-series variants.

    Machine learning methodsIntermediate3 min
  • Data leakage in machine learning: how good scores lie

    Data leakage lets outside information slip into training, inflating scores that collapse in production. Learn the common types and how to prevent them.

    Machine learning methodsIntermediate3 min
  • How to chunk documents for RAG

    Chunking decides what your RAG system can retrieve. Compare fixed-size, recursive, heading-aware and semantic chunking, and learn how to pick size and overlap.

    Retrieval-augmented generationIntermediate3 min
  • How to evaluate a RAG system: retrieval and answer quality

    Measure a RAG system in two layers: did retrieval find the right passages, and is the answer faithful to them? Learn recall@k, MRR, faithfulness and a practical test set.

    Retrieval-augmented generationIntermediate3 min
  • How vector search works: similarity, top-k and hybrid search

    Vector search finds text by meaning instead of exact words. Learn cosine similarity, top-k retrieval, approximate nearest neighbour indexes, BM25 and hybrid ranking.

    Retrieval-augmented generationIntermediate3 min
  • HTTP security headers explained: CSP, HSTS and more

    Security headers tell browsers how to protect your users. Learn what CSP, HSTS, X-Content-Type-Options, Referrer-Policy, Permissions-Policy and frame-ancestors do.

    Web securityBeginner3 min
  • Overfitting and underfitting in machine learning

    Overfitting means a model memorises its training data; underfitting means it misses the pattern. Learn to spot both from your scores and how to fix each.

    Machine learning methodsBeginner3 min
  • Precision, recall and F1 score explained with an example

    Accuracy can be misleading on imbalanced data. Learn the confusion matrix, precision, recall, F1 and thresholds through a worked phishing-detection example.

    Machine learning methodsBeginner3 min
  • Prompt injection in RAG systems and how to reduce it

    Prompt injection makes a language model follow attacker-written text. Learn direct and indirect injection in RAG, why filters fail, and layered defences.

    AI securityIntermediate3 min
  • TCP vs UDP: differences, handshakes and when each is used

    TCP gives reliable, ordered delivery after a handshake; UDP sends independent datagrams. Compare headers, use cases, ports and how both look in Wireshark.

    NetworkingBeginner3 min
  • Train, validation and test sets explained

    Why machine learning data is split three ways, what each split is for, how to split correctly with scikit-learn, and the mistakes that make a test score meaningless.

    Machine learning methodsBeginner3 min
  • What are embeddings? Meaning as numbers, explained

    Embeddings turn text into lists of numbers so that similar meanings land close together. Learn how they work, how similarity is measured, and where they fail.

    LLM foundationsBeginner3 min
  • What is a context window in an LLM?

    The context window is the maximum number of tokens a language model can consider at once. Learn what fills it, what happens when it overflows, and how to budget it.

    LLM foundationsBeginner3 min
  • What is a token in an LLM?

    A token is the unit of text a language model reads and writes. Learn how tokenizers split text, why token counts differ by language, and why it matters for cost.

    LLM foundationsBeginner4 min
  • What is retrieval-augmented generation (RAG)?

    RAG makes a language model answer from your own documents by retrieving relevant passages first. Learn the pipeline step by step, and when RAG is the right choice.

    Retrieval-augmented generationBeginner3 min
  • Why long context windows still miss things

    A million-token context window does not mean a model uses every token well. Learn about the lost-in-the-middle effect, attention cost, and how to test it yourself.

    LLM foundationsIntermediate3 min