Guided series

RAG Systems from Scratch

Build a retrieval-augmented generation pipeline, then make it genuinely good: chunking, retrieval tricks, evaluation, caching and scale.

11 articles · in reading order

  1. Building a RAG Pipeline from Scratch with Python (2026)

    Build a complete RAG pipeline from scratch with Python. No frameworks, no black boxes — just plain Python code you actually understand.

    September 1, 2026
  2. Evaluating RAG Systems: Metrics That Actually Matter

    Learn how to evaluate RAG systems with the right metrics. Precision recall faithfulness and relevance benchmarks that actually improve your pipeline.

    September 1, 2026
  3. Query Expansion and Rewriting for Better RAG Retrieval

    Query expansion and rewriting for RAG: LLM query rewriting, MultiQueryRetriever, HyDE, query decomposition, and step-back prompting — measured with RAGAS.

    September 24, 2026
  4. HyDE Explained: Hypothetical Document Embeddings for RAG

    HyDE explained: Hypothetical Document Embeddings close the query-document gap with LangChain HypotheticalDocumentEmbedder, custom prompts, LlamaIndex transforms, and RAGAS measurement.

    September 23, 2026
  5. Parent-Child Chunking: Advanced Document Splitting for RAG

    Parent-child chunking tutorial for RAG: small-to-big retrieval with LlamaIndex HierarchicalNodeParser, AutoMergingRetriever, and LangChain ParentDocumentRetriever.

    September 24, 2026
  6. Contextual Compression in RAG: Retrieve Less, Answer Better

    Contextual compression RAG tutorial with LangChain: LLMChainExtractor, EmbeddingsFilter, reranker pipelines, and measuring gains with RAGAS context precision.

    September 24, 2026
  7. Agentic RAG: Combining AI Agents with Retrieval

    Agentic RAG tutorial: build grade-and-retry retrieval loops with LangGraph, Corrective RAG (CRAG), LlamaIndex ReActAgent, and production guardrails for cost and latency.

    September 24, 2026
  8. GraphRAG Explained: Knowledge Graphs Meet Retrieval-Augmented Generation

    GraphRAG explained: knowledge graphs for multi-hop and global RAG queries with Microsoft GraphRAG, Neo4j + LlamaIndex, community summarization, and hybrid agentic routing.

    September 24, 2026
  9. Caching Strategies for RAG: Reduce Latency and API Costs

    Four RAG caching layers explained: embedding cache, retrieval cache, semantic response cache, and prompt caching, with threshold tuning and hit-rate monitoring.

    September 25, 2026
  10. Scaling RAG to Millions of Documents: Architecture Patterns

    Six architecture patterns for scaling RAG to millions of documents: hybrid retrieval, chunking, reranking, index sharding, incremental ingestion, and caching.

    September 25, 2026
  11. RAGAS Tutorial: Automated Evaluation for RAG Pipelines

    RAGAS tutorial: evaluate RAG pipelines with faithfulness, answer relevancy, context precision, and context recall metrics using LLM-as-a-judge and synthetic testsets.

    September 24, 2026

Want more paths? Browse all series or pick a topic.

What are You Looking For?

esc