Skip to content

06. What is Retrieval-Augmented Generation (RAG)?

Retrieval-Augmented Generation (RAG) is the architecture that lets an LLM answer questions about documents it has never seen — by retrieving relevant information and feeding it into the prompt at query time.

If you’ve ever uploaded a PDF to ChatGPT and asked questions about it, you’ve used RAG. If you’ve used NotebookLM to ask about your notes, Perplexity to search with citations, or Cursor to understand your codebase — that’s RAG. It is the single most important pattern in production AI engineering.


Imagine a doctor who has years of medical training. She knows anatomy, pharmacology, and diagnosis. But before treating a new patient, she doesn’t rely on memory alone. She reads the patient’s chart — latest blood work, medical history, current symptoms.

The doctor’s training is the LLM. The patient’s chart is the retrieved knowledge.

Without the chart, the doctor can only give general advice. With the chart, she gives a precise, personalized answer.

That’s RAG.

flowchart LR
subgraph WITHOUT_RAG["Without RAG"]
A1["User: 'What did we decide\nin yesterday's meeting?'"] --> LLM["LLM\n(only knows training data)"]
LLM --> A2["❌ 'I don't have\naccess to that information'"]
end
subgraph WITH_RAG["With RAG"]
B1["User: 'What did we decide\nin yesterday's meeting?'"] --> RET["🔍 Retriever\n(finds relevant notes)"]
B2["📁 Meeting Notes\nDatabase"] --> RET
RET --> AUG["📝 Augmented Prompt\n(question + meeting notes)"]
AUG --> LLM2["LLM\n(reads + answers)"]
LLM2 --> B3["✅ 'We decided to launch\nthe feature next quarter'"]
end
style WITHOUT_RAG fill:#ef4444,color:#fff
style WITH_RAG fill:#22c55e,color:#fff

In school, there are two types of exams:

Closed-book: You must answer from memory. If you didn’t study it, you can’t answer it. This is a pure LLM — limited to its training data.

Open-book: You can refer to your notes and textbook during the exam. The questions test your ability to find and apply information, not just memorize. This is RAG.

RAG turns every LLM interaction into an open-book exam. The model doesn’t need to memorize everything — it just needs to know where to look.


The system searches a knowledge base (vector database) for documents relevant to the user’s question. This involves embedding the question and finding nearest neighbors.

The retrieved documents are inserted into the LLM’s prompt alongside the original question. The prompt tells the LLM: “Answer based on these documents.”

The LLM generates a response using the retrieved context. Because the relevant information is right there in the prompt, the answer is accurate and grounded in your data.

flowchart TD
subgraph RETRIEVAL["🔄 Stage 1: Retrieval"]
R1["User Question"] --> R2["Embedding Model"]
R2 --> R3["Similarity Search\n(Vector Database)"]
R3 --> R4["Top K Relevant\nDocument Chunks"]
end
subgraph AUGMENT["📝 Stage 2: Augmentation"]
R4 --> A1["Prompt Builder"]
A1 --> A2["System: Answer using the context below.\n\nContext: [retrieved chunks]\n\nUser: [original question]"]
end
subgraph GENERATION["💬 Stage 3: Generation"]
A2 --> G1["LLM\n(reads prompt + context)"]
G1 --> G2["Accurate Answer\n(grounded in your data)"]
end
style RETRIEVAL fill:#3b82f6,color:#fff
style AUGMENT fill:#8b5cf6,color:#fff
style GENERATION fill:#22c55e,color:#fff

sequenceDiagram
participant User
participant App as Your App
participant Retriever
participant VDB as Vector DB
participant LLM
User->>App: "What was our Q3 revenue?"
App->>Retriever: Embed question
Retriever->>VDB: Find similar vectors
VDB-->>Retriever: Top 5 chunks
Retriever-->>App: Relevant document pieces
App->>LLM: Prompt + retrieved context
LLM-->>App: "Q3 revenue was $12.4M..."
App-->>User: ✅ Answer with citation

ProblemWithout RAGWith RAG
Knowledge cutoffLLM doesn’t know recent eventsRetrieves current documents
Private dataLLM has never seen your docsRetrieves from your knowledge base
HallucinationsLLM guesses when unsureGrounds answers in retrieved facts
CostFine-tuning is expensiveNo training needed, just indexing
UpdatesMust retrain to add knowledgeJust add new documents
ApproachBest ForCostUpdate Frequency
RAGSpecific knowledge, private data, frequent updatesLow (per-query)Instant
Fine-tuningChanging model behavior, tone, formatHigh (training)Weeks
Prompt engineeringSimple instructions, no new knowledgeZeroInstant
RetrainingUpdating core knowledgeVery highMonths
flowchart TD
Q["Do you need the model to
know specific, private,
or recent information?"] -->|"Yes"| RAG["Use RAG\n(cheap, instant updates)"]
Q -->|"No"| Q2["Do you need to change
the model's behavior,
tone, or output format?"]
Q2 -->|"Yes"| FT["Use Fine-Tuning\n(expensive, behavioral change)"]
Q2 -->|"No"| PE["Use Prompt Engineering\n(zero cost, simple instructions)"]
RAG --> COMBINE["Often used together!"]
FT --> COMBINE
style RAG fill:#22c55e,color:#fff
style FT fill:#f59e0b,color:#fff
style PE fill:#3b82f6,color:#fff

ProductWhat It RetrievesHow RAG Helps
ChatPDFYour uploaded PDFsAsk questions about any document
NotebookLMYour notes, sources, transcriptsResearch assistant that knows your materials
Cursor / CopilotYour codebaseCode completion aware of your project
PerplexityWeb search resultsAnswers with real-time sources
Claude ProjectsYour uploaded filesProject-specific knowledge
Customer Support BotsKnowledge base articlesAccurate answers from company docs

MistakeWhy It’s Wrong
❌ “RAG replaces fine-tuning”They serve different purposes. RAG adds knowledge; fine-tuning changes behavior. Use both together
❌ “I can send my entire knowledge base in one prompt”Context windows are limited. RAG selects the most relevant pieces — you can’t fit everything
❌ “RAG guarantees no hallucinations”RAG reduces hallucinations by grounding answers, but the LLM can still ignore or misinterpret retrieved context

Q: What does RAG stand for and what does it do?

Retrieval-Augmented Generation. It retrieves relevant documents from a knowledge base, augments the LLM’s prompt with them, and generates a response grounded in those documents.

Q: What are the three stages of RAG and what happens in each?

(1) Retrieval — search a vector database for documents similar to the query. (2) Augmentation — insert retrieved documents into the LLM prompt. (3) Generation — LLM reads the prompt and generates an answer based on the provided context.

Q: Design a RAG system for a customer support chatbot that serves 10,000 users. Consider latency, cost, and accuracy.

Architecture: (1) Indexing pipeline — chunk support docs, embed with text-embedding-3-small, store in Pinecone/Qdrant with metadata filters (product, category). (2) Query pipeline — user question → embed → hybrid search (semantic + keyword) → top 5 chunks → rerank with cross-encoder → LLM generates answer. (3) Caching — cache frequent queries (same question, same answer). (4) Monitoring — track retrieval precision, hallucination rate, user feedback. (5) Cost optimization — use smaller LLM (GPT-4o-mini) for simple queries, fall back to larger model for complex ones.


ConceptKey Point
RAGArchitecture that grounds LLM answers in retrieved documents
RetrievalFind relevant documents from a knowledge base
AugmentationInsert retrieved context into the LLM prompt
GenerationLLM produces answer using the provided context
Why it mattersAccurate, up-to-date, private, no retraining needed

Previous: 05 — Vector Databases Introduction →

Next: 07 — Document Chunking →