Guided series
RAG Systems from Scratch
Build a retrieval-augmented generation pipeline, then make it genuinely good: chunking, retrieval tricks, evaluation, caching and scale.
11 articles · in reading order
-
Building a RAG Pipeline from Scratch with Python (2026)
Build a complete RAG pipeline from scratch with Python. No frameworks, no black boxes — just plain Python code you actually understand.
September 1, 2026 -
Evaluating RAG Systems: Metrics That Actually Matter
Learn how to evaluate RAG systems with the right metrics. Precision recall faithfulness and relevance benchmarks that actually improve your pipeline.
September 1, 2026 -
Query Expansion and Rewriting for Better RAG Retrieval
Query expansion and rewriting for RAG: LLM query rewriting, MultiQueryRetriever, HyDE, query decomposition, and step-back prompting — measured with RAGAS.
September 24, 2026 -
HyDE Explained: Hypothetical Document Embeddings for RAG
HyDE explained: Hypothetical Document Embeddings close the query-document gap with LangChain HypotheticalDocumentEmbedder, custom prompts, LlamaIndex transforms, and RAGAS measurement.
September 23, 2026 -
Parent-Child Chunking: Advanced Document Splitting for RAG
Parent-child chunking tutorial for RAG: small-to-big retrieval with LlamaIndex HierarchicalNodeParser, AutoMergingRetriever, and LangChain ParentDocumentRetriever.
September 24, 2026 -
Contextual Compression in RAG: Retrieve Less, Answer Better
Contextual compression RAG tutorial with LangChain: LLMChainExtractor, EmbeddingsFilter, reranker pipelines, and measuring gains with RAGAS context precision.
September 24, 2026 -
Agentic RAG: Combining AI Agents with Retrieval
Agentic RAG tutorial: build grade-and-retry retrieval loops with LangGraph, Corrective RAG (CRAG), LlamaIndex ReActAgent, and production guardrails for cost and latency.
September 24, 2026 -
GraphRAG Explained: Knowledge Graphs Meet Retrieval-Augmented Generation
GraphRAG explained: knowledge graphs for multi-hop and global RAG queries with Microsoft GraphRAG, Neo4j + LlamaIndex, community summarization, and hybrid agentic routing.
September 24, 2026 -
Caching Strategies for RAG: Reduce Latency and API Costs
Four RAG caching layers explained: embedding cache, retrieval cache, semantic response cache, and prompt caching, with threshold tuning and hit-rate monitoring.
September 25, 2026 -
Scaling RAG to Millions of Documents: Architecture Patterns
Six architecture patterns for scaling RAG to millions of documents: hybrid retrieval, chunking, reranking, index sharding, incremental ingestion, and caching.
September 25, 2026 -
RAGAS Tutorial: Automated Evaluation for RAG Pipelines
RAGAS tutorial: evaluate RAG pipelines with faithfulness, answer relevancy, context precision, and context recall metrics using LLM-as-a-judge and synthetic testsets.
September 24, 2026
Want more paths? Browse all series or pick a topic.