Learn how AI reads, retrieves and breaks.
Plain-language lessons on tokens, context windows, retrieval-augmented generation, machine learning methods and the security problems they create. With labs you run yourself, browser tools, and a tutor that shows where every answer came from.
[29387, 1495, 20837, 31523, 11, 171243, 326, 29617, 13]
- Tokens
- 9
- Characters
- 41
Counted with the o200k_base tokenizer. Try your own text
What you can do here
- Read16 articles
- Explanations that start with a direct answer, then show the mechanism.
- Build3 labs
- Step-by-step exercises on your own machine, from a tiny RAG system to Wireshark.
- Try8 tools
- Token counter, chunking playground, subnet calculator and more. All run in your browser.
- AskAI tutor
- Questions answered from this site's material, with the retrieved sources shown.
Learning paths
All pathsOrdered lessons with a reason for each step. Pick one and follow it through.
- Artificial intelligenceBeginner
How language models read text
Tokens, context windows and embeddings: what a model actually receives.
4 articles
- Artificial intelligenceIntermediate
Build and evaluate RAG systems
Chunking, retrieval, grounded prompts, evaluation and security for answering from your own documents.
5 articles and 1 lab
- Artificial intelligenceBeginner
Machine learning methods that hold up
Data splits, overfitting, cross-validation, the right metrics and data leakage.
5 articles and 1 lab
- SecurityBeginner
Network and web security basics
Transport protocols, packet captures and the browser protections every site should enable.
2 articles and 1 lab
A study tutor that shows its sources
Ask a question and the tutor first searches this site's articles and labs. It answers only from the passages it finds, cites them, and says so when the material does not cover your question.
Next to every answer you can see the retrieved passages, their scores and how many tokens they used. The tutor is also a working example of the ideas in the RAG path.
Why do models miss facts in the middle of a long prompt?
- [1]Why long context windows still miss things > Short answer16.91
- [2]Why long context windows still miss things > The lost-in-the-middle finding15.11
- [3]Why long context windows still miss things > FAQ13.53
3 passages, 366 tokens of context. Scores are BM25 keyword relevance.
Recent articles
All articlesCross-validation explained: k-fold, stratified and grouped
Cross-validation gives a more reliable performance estimate than one split by rotating the validation fold. Learn k-fold, stratified, group and time-series variants.
Machine learning methodsIntermediate3 minData leakage in machine learning: how good scores lie
Data leakage lets outside information slip into training, inflating scores that collapse in production. Learn the common types and how to prevent them.
Machine learning methodsIntermediate3 minHow to chunk documents for RAG
Chunking decides what your RAG system can retrieve. Compare fixed-size, recursive, heading-aware and semantic chunking, and learn how to pick size and overlap.
Retrieval-augmented generationIntermediate3 minHow to evaluate a RAG system: retrieval and answer quality
Measure a RAG system in two layers: did retrieval find the right passages, and is the answer faithful to them? Learn recall@k, MRR, faithfulness and a practical test set.
Retrieval-augmented generationIntermediate3 minHow vector search works: similarity, top-k and hybrid search
Vector search finds text by meaning instead of exact words. Learn cosine similarity, top-k retrieval, approximate nearest neighbour indexes, BM25 and hybrid ranking.
Retrieval-augmented generationIntermediate3 minHTTP security headers explained: CSP, HSTS and more
Security headers tell browsers how to protect your users. Learn what CSP, HSTS, X-Content-Type-Options, Referrer-Policy, Permissions-Policy and frame-ancestors do.
Web securityBeginner3 min
Two subjects, one overlap
AI systems are software on networks, and they bring new ways to fail. Where the two meet, you get AI security.
Artificial intelligence
How machine learning is done properly, how language models read text, and how retrieval-augmented generation (RAG) systems are built and evaluated.
Cybersecurity
Networking and security fundamentals, and the new attack surface that AI systems add: prompt injection, poisoned documents and data leakage.
Labs
All labs- Measure overfitting with validation curves
Beginner30 minutes
- Watch a TCP handshake in Wireshark
Beginner30 to 40 minutes
- Build a tiny RAG system in Python
Intermediate45 to 60 minutes
Tools
All tools- Token counter
Count tokens with a real tokenizer and see how much of a context window your text fills.
- Chunking playground
Split your own text with different RAG chunking strategies and compare the pieces.
- Prompt injection checker
See which common injection patterns a simple filter spots, and learn why filters are not enough.
- IPv4 subnet calculator
Network, broadcast, host range and mask for any IPv4 address in CIDR notation.
- Security headers analyzer
Paste a site's response headers and get a checklist of missing or weak protections.