Skip to content

01. Why Retrieval Systems Exist

A Large Language Model knows what the internet knew when it stopped training. A Retrieval System knows what you asked it to read five seconds ago.

This is the single most important concept in production AI: LLMs don’t automatically know your data. You have to give it to them. Retrieval Systems are how you do that.


ChatGPT was trained on data up to a certain date. Ask it about your company’s Q3 earnings report (released yesterday) and it has no idea. Ask it about a confidential legal document you just uploaded — it can’t see it. Ask it about the internal API documentation your team wrote last week — nothing.

This isn’t a flaw. It’s a fundamental constraint:

ConceptWhat It Means
Training DataEverything the model learned during training
Knowledge CutoffThe date when training ended
Private DataDocuments the model has never seen
RetrievalGiving the model relevant information at query time

Imagine a brilliant student who graduated top of her class. She knows everything from her textbooks. But tomorrow, the curriculum changes. New research is published. Company policies are updated.

If you hand her the new textbook right before the exam, she can read it and answer questions about it.

That’s retrieval.

Without the textbook, she can only answer based on what she already knows — which might be outdated or incomplete.


You walk into a library. The librarian has read every book in existence — but all the books are from 2022. You need information from a research paper published last month.

The librarian can’t magically know the paper. But if you hand them the paper and ask a question, they can read it and answer.

That’s the difference between a pure LLM and an LLM with retrieval.

flowchart LR
subgraph WITHOUT["Without Retrieval"]
A1["User Question"] --> LLM["LLM\n(knowledge cutoff only)"]
LLM --> A2["Answer based on\nold training data"]
end
subgraph WITH["With Retrieval"]
B1["User Question"] --> RET["Retrieval System\n(finds relevant docs)"]
B2["Your Documents\n(PDFs, wikis, DBs)"] --> RET
RET --> B3["LLM\n(prompt + retrieved context)"]
B3 --> B4["Answer based on\nYOUR data"]
end
style WITHOUT fill:#ef4444,color:#fff
style WITH fill:#22c55e,color:#fff

When an LLM is trained, billions of parameters are updated to predict the next token. This process takes months and costs millions of dollars. You cannot (and should not) retrain a model every time a new document is created.

Every LLM has a training cutoff date:

ModelKnowledge Cutoff
GPT-4o~Oct 2023
Claude 3.5~Apr 2024
Gemini 1.5~Mar 2023
Llama 3~Dec 2023

Anything after that date is unknown to the model unless you provide it at query time.

Your company’s internal documents, your personal notes, your confidential contracts — none of these exist in any public training dataset. The model has literally never seen them.

flowchart TD
subgraph CAN_KNOW["What the LLM Knows"]
A["Public Internet Data\n(up to cutoff date)"]
B["Common Knowledge\n(science, history, literature)"]
C["Programming Languages\n(public code repositories)"]
end
subgraph CANNOT_KNOW["What the LLM Does NOT Know"]
D["Your Company Documents"]
E["Real-Time Data\n(today's news, stock prices)"]
F["Private Conversations"]
G["Post-Cutoff Events"]
end
CAN_KNOW --> H["LLM"]
CANNOT_KNOW -.->|"❌ Cannot access"| H
CANNOT_KNOW -->|"✅ Retrieved at query time"| I["Augmented LLM"]
style CAN_KNOW fill:#22c55e,color:#fff
style CANNOT_KNOW fill:#ef4444,color:#fff
style H fill:#f59e0b,color:#fff
style I fill:#3b82f6,color:#fff

flowchart LR
subgraph TRAINING["Training"]
T1["Trillions of tokens\n(months of GPU time)"] --> T2["Model Weights Updated\n(costs millions $)"]
T2 --> T3["Knowledge is\npermanent but frozen"]
end
subgraph RETRIEVAL["Retrieval"]
R1["Upload a PDF\n(5 seconds)"] --> R2["Document is\nindexed and stored"]
R2 --> R3["Knowledge is\navailable immediately"]
end
style TRAINING fill:#f59e0b,color:#fff
style RETRIEVAL fill:#3b82f6,color:#fff
AspectTrainingRetrieval
CostMillions of dollarsPennies per query
TimeWeeks to monthsMilliseconds
FreshnessFrozen at cutoffAlways up-to-date
PrivacyPublic data onlyYour private data
UpdateRetrain from scratchAdd new documents instantly

Retrieval is the process of finding relevant information from a knowledge base and providing it to the LLM as part of the prompt.

User: "What was our Q3 revenue?"
Step 1 — Retrieve: Search the knowledge base for documents about Q3 revenue
Step 2 — Augment: Insert those documents into the prompt
Step 3 — Generate: LLM reads the documents and answers
Prompt (simplified):
System: Answer using the context below.
Context: [Q3 revenue was $12.4M, up 18% YoY...]
User: What was our Q3 revenue?
Response: Your Q3 revenue was $12.4M, up 18% year over year.

ProductWhat It RetrievesWhy It Matters
ChatGPT with SearchLive web resultsAnswers about current events
PerplexityWeb pages + citationsEvery answer has sources
NotebookLMYour uploaded documentsAnswer questions about your notes
GitHub CopilotYour codebaseAutocomplete uses your code context
Claude ProjectsYour uploaded filesProject-specific knowledge
CursorYour codebaseAI that knows your entire project

MistakeWhy It’s Wrong
❌ “I’ll just retrain the model on my documents”Impractical for most use cases — expensive, slow, and the model may forget old knowledge (catastrophic forgetting)
❌ “I’ll put everything in the prompt”Context windows are limited (128K-200K tokens). You can’t fit an entire company wiki in one prompt
❌ “The model will remember what I told it earlier”LLMs have no persistent memory between conversations (unless you implement it via retrieval)

Q: Why can’t an LLM answer questions about documents it wasn’t trained on?

Because the model’s knowledge is frozen at the time of training. It doesn’t have access to new or private documents unless they are provided at query time through retrieval.

Q: What’s the difference between training a model on your data and retrieving documents at query time?

Training updates the model’s weights — expensive, slow, permanent. Retrieval finds relevant documents at query time and inserts them into the prompt — cheap, fast, and easy to update. Use training for fundamental capabilities, retrieval for specific knowledge.

Q: Design a retrieval system for a company with 10,000 internal documents. What are the key considerations?

Key considerations: (1) Chunking strategy — how to split documents into searchable pieces, (2) Embedding model choice — balancing accuracy vs cost, (3) Vector database — scalability and filtering, (4) Hybrid search — combining keyword and semantic search for better results, (5) Relevance scoring — ensuring only high-quality results reach the LLM, (6) Caching — reducing latency for frequent queries, (7) Access control — ensuring users only retrieve documents they’re authorized to see.


ConceptKey Point
Knowledge CutoffLLMs only know data up to their training date
Private DataLLMs cannot see your documents unless you provide them
RetrievalFinding relevant documents at query time
AugmentationInserting retrieved documents into the LLM’s prompt
Why It MattersWithout retrieval, LLMs are limited to public, frozen knowledge

Previous: Phase 4: Large Language Models →

Next: 02 — Embeddings Deep Dive →