06. What is Retrieval-Augmented Generation (RAG)?
Introduction
Section titled “Introduction”Retrieval-Augmented Generation (RAG) is the architecture that lets an LLM answer questions about documents it has never seen — by retrieving relevant information and feeding it into the prompt at query time.
If you’ve ever uploaded a PDF to ChatGPT and asked questions about it, you’ve used RAG. If you’ve used NotebookLM to ask about your notes, Perplexity to search with citations, or Cursor to understand your codebase — that’s RAG. It is the single most important pattern in production AI engineering.
Why This Concept Exists
Section titled “Why This Concept Exists”The Story
Section titled “The Story”Imagine a doctor who has years of medical training. She knows anatomy, pharmacology, and diagnosis. But before treating a new patient, she doesn’t rely on memory alone. She reads the patient’s chart — latest blood work, medical history, current symptoms.
The doctor’s training is the LLM. The patient’s chart is the retrieved knowledge.
Without the chart, the doctor can only give general advice. With the chart, she gives a precise, personalized answer.
That’s RAG.
flowchart LR subgraph WITHOUT_RAG["Without RAG"] A1["User: 'What did we decide\nin yesterday's meeting?'"] --> LLM["LLM\n(only knows training data)"] LLM --> A2["❌ 'I don't have\naccess to that information'"] end
subgraph WITH_RAG["With RAG"] B1["User: 'What did we decide\nin yesterday's meeting?'"] --> RET["🔍 Retriever\n(finds relevant notes)"] B2["📁 Meeting Notes\nDatabase"] --> RET RET --> AUG["📝 Augmented Prompt\n(question + meeting notes)"] AUG --> LLM2["LLM\n(reads + answers)"] LLM2 --> B3["✅ 'We decided to launch\nthe feature next quarter'"] end
style WITHOUT_RAG fill:#ef4444,color:#fff style WITH_RAG fill:#22c55e,color:#fffReal-World Analogy
Section titled “Real-World Analogy”The Open-Book Exam
Section titled “The Open-Book Exam”In school, there are two types of exams:
Closed-book: You must answer from memory. If you didn’t study it, you can’t answer it. This is a pure LLM — limited to its training data.
Open-book: You can refer to your notes and textbook during the exam. The questions test your ability to find and apply information, not just memorize. This is RAG.
RAG turns every LLM interaction into an open-book exam. The model doesn’t need to memorize everything — it just needs to know where to look.
The Three Stages of RAG
Section titled “The Three Stages of RAG”1. Retrieval
Section titled “1. Retrieval”The system searches a knowledge base (vector database) for documents relevant to the user’s question. This involves embedding the question and finding nearest neighbors.
2. Augmentation
Section titled “2. Augmentation”The retrieved documents are inserted into the LLM’s prompt alongside the original question. The prompt tells the LLM: “Answer based on these documents.”
3. Generation
Section titled “3. Generation”The LLM generates a response using the retrieved context. Because the relevant information is right there in the prompt, the answer is accurate and grounded in your data.
flowchart TD subgraph RETRIEVAL["🔄 Stage 1: Retrieval"] R1["User Question"] --> R2["Embedding Model"] R2 --> R3["Similarity Search\n(Vector Database)"] R3 --> R4["Top K Relevant\nDocument Chunks"] end
subgraph AUGMENT["📝 Stage 2: Augmentation"] R4 --> A1["Prompt Builder"] A1 --> A2["System: Answer using the context below.\n\nContext: [retrieved chunks]\n\nUser: [original question]"] end
subgraph GENERATION["💬 Stage 3: Generation"] A2 --> G1["LLM\n(reads prompt + context)"] G1 --> G2["Accurate Answer\n(grounded in your data)"] end
style RETRIEVAL fill:#3b82f6,color:#fff style AUGMENT fill:#8b5cf6,color:#fff style GENERATION fill:#22c55e,color:#fffThe Complete RAG Flow
Section titled “The Complete RAG Flow”sequenceDiagram participant User participant App as Your App participant Retriever participant VDB as Vector DB participant LLM
User->>App: "What was our Q3 revenue?" App->>Retriever: Embed question Retriever->>VDB: Find similar vectors VDB-->>Retriever: Top 5 chunks Retriever-->>App: Relevant document pieces App->>LLM: Prompt + retrieved context LLM-->>App: "Q3 revenue was $12.4M..." App-->>User: ✅ Answer with citationWhy RAG Exists
Section titled “Why RAG Exists”The Problem RAG Solves
Section titled “The Problem RAG Solves”| Problem | Without RAG | With RAG |
|---|---|---|
| Knowledge cutoff | LLM doesn’t know recent events | Retrieves current documents |
| Private data | LLM has never seen your docs | Retrieves from your knowledge base |
| Hallucinations | LLM guesses when unsure | Grounds answers in retrieved facts |
| Cost | Fine-tuning is expensive | No training needed, just indexing |
| Updates | Must retrain to add knowledge | Just add new documents |
When to Use RAG vs Alternatives
Section titled “When to Use RAG vs Alternatives”| Approach | Best For | Cost | Update Frequency |
|---|---|---|---|
| RAG | Specific knowledge, private data, frequent updates | Low (per-query) | Instant |
| Fine-tuning | Changing model behavior, tone, format | High (training) | Weeks |
| Prompt engineering | Simple instructions, no new knowledge | Zero | Instant |
| Retraining | Updating core knowledge | Very high | Months |
RAG vs Fine-Tuning: Decision Flow
Section titled “RAG vs Fine-Tuning: Decision Flow”flowchart TD Q["Do you need the model toknow specific, private,or recent information?"] -->|"Yes"| RAG["Use RAG\n(cheap, instant updates)"] Q -->|"No"| Q2["Do you need to changethe model's behavior,tone, or output format?"] Q2 -->|"Yes"| FT["Use Fine-Tuning\n(expensive, behavioral change)"] Q2 -->|"No"| PE["Use Prompt Engineering\n(zero cost, simple instructions)"] RAG --> COMBINE["Often used together!"] FT --> COMBINE
style RAG fill:#22c55e,color:#fff style FT fill:#f59e0b,color:#fff style PE fill:#3b82f6,color:#fffReal Production Examples
Section titled “Real Production Examples”| Product | What It Retrieves | How RAG Helps |
|---|---|---|
| ChatPDF | Your uploaded PDFs | Ask questions about any document |
| NotebookLM | Your notes, sources, transcripts | Research assistant that knows your materials |
| Cursor / Copilot | Your codebase | Code completion aware of your project |
| Perplexity | Web search results | Answers with real-time sources |
| Claude Projects | Your uploaded files | Project-specific knowledge |
| Customer Support Bots | Knowledge base articles | Accurate answers from company docs |
Common Mistakes
Section titled “Common Mistakes”| Mistake | Why It’s Wrong |
|---|---|
| ❌ “RAG replaces fine-tuning” | They serve different purposes. RAG adds knowledge; fine-tuning changes behavior. Use both together |
| ❌ “I can send my entire knowledge base in one prompt” | Context windows are limited. RAG selects the most relevant pieces — you can’t fit everything |
| ❌ “RAG guarantees no hallucinations” | RAG reduces hallucinations by grounding answers, but the LLM can still ignore or misinterpret retrieved context |
Interview Questions
Section titled “Interview Questions”Q: What does RAG stand for and what does it do?
Retrieval-Augmented Generation. It retrieves relevant documents from a knowledge base, augments the LLM’s prompt with them, and generates a response grounded in those documents.
Intermediate
Section titled “Intermediate”Q: What are the three stages of RAG and what happens in each?
(1) Retrieval — search a vector database for documents similar to the query. (2) Augmentation — insert retrieved documents into the LLM prompt. (3) Generation — LLM reads the prompt and generates an answer based on the provided context.
Senior - Architecture
Section titled “Senior - Architecture”Q: Design a RAG system for a customer support chatbot that serves 10,000 users. Consider latency, cost, and accuracy.
Architecture: (1) Indexing pipeline — chunk support docs, embed with text-embedding-3-small, store in Pinecone/Qdrant with metadata filters (product, category). (2) Query pipeline — user question → embed → hybrid search (semantic + keyword) → top 5 chunks → rerank with cross-encoder → LLM generates answer. (3) Caching — cache frequent queries (same question, same answer). (4) Monitoring — track retrieval precision, hallucination rate, user feedback. (5) Cost optimization — use smaller LLM (GPT-4o-mini) for simple queries, fall back to larger model for complex ones.
Summary
Section titled “Summary”| Concept | Key Point |
|---|---|
| RAG | Architecture that grounds LLM answers in retrieved documents |
| Retrieval | Find relevant documents from a knowledge base |
| Augmentation | Insert retrieved context into the LLM prompt |
| Generation | LLM produces answer using the provided context |
| Why it matters | Accurate, up-to-date, private, no retraining needed |
Navigation
Section titled “Navigation”Previous: 05 — Vector Databases Introduction →
Next: 07 — Document Chunking →