Skip to content

25. Phase Summary & Project Roadmap

Welcome to the end of Phase 5. You’ve built ChatPDF, a Company Knowledge Assistant, a Documentation Chatbot, and a GitHub Code Assistant. Now it’s time to review everything, compare approaches, and chart your path forward.

This summary page consolidates everything you’ve learned across 24 documents and 5 chunks. Use it as a reference, a decision guide, and a roadmap to further mastery.


flowchart TD
subgraph DOCUMENTS["Document Sources"]
PDF["📄 PDFs"]
WEB["🌐 Websites"]
REPO["💻 Code Repos"]
WIKI["📝 Internal Wikis"]
DB["🗄️ Databases"]
end
subgraph INGESTION["Ingestion Pipeline"]
PARSE["Parse & Extract"]
CHUNK["Chunk Documents"]
EMBED["Generate Embeddings"]
STORE["Store in Vector DB"]
end
subgraph RETRIEVAL["Retrieval Strategies"]
SIMPLE["Basic Vector Search"]
HYBRID["Hybrid Search\n(BM25 + Vector)"]
MQ["Multi-Query\n(Query Expansion)"]
PC["Parent-Child\n(Chunk Expansion)"]
RERANK["Re-ranking\n(Cross-Encoder)"]
end
subgraph GEN["Generation"]
PROMPT["Prompt Builder"]
COMPRESS["Context Compression"]
LLM["LLM Generation"]
CITATION["Citation Builder"]
end
subgraph PROD["Production"]
CACHE["Caching Layer"]
AUTH["Auth & RBAC"]
MONITOR["Monitoring & Evaluation"]
SCALE["Scaling & Load Balancing"]
end
DOCUMENTS --> INGESTION --> RETRIEVAL --> GEN --> PROD
style DOCUMENTS fill:#3b82f6,color:#fff
style INGESTION fill:#8b5cf6,color:#fff
style RETRIEVAL fill:#f59e0b,color:#fff
style GEN fill:#22c55e,color:#fff
style PROD fill:#ef4444,color:#fff

flowchart TD
subgraph EVOLUTION["RAG Evolution"]
BASIC["Basic RAG\nEmbed → Search → Generate"]
ADV["Advanced RAG\nHybrid + Re-rank + Compression"]
ENTERPRISE["Enterprise RAG\nMulti-tenant + Security + Eval"]
AGENTIC["Agentic RAG\n(Covered in Phase 6)\nSelf-query + Tool use + Planning"]
end
BASIC --> ADV --> ENTERPRISE --> AGENTIC
style BASIC fill:#3b82f6,color:#fff
style ADV fill:#8b5cf6,color:#fff
style ENTERPRISE fill:#22c55e,color:#fff
style AGENTIC fill:#ef4444,color:#fff
DimensionBasic RAGAdvanced RAGEnterprise RAG
SearchVector onlyHybrid (vector + keyword)Hybrid + multi-query + parent-child
RankingVector similarity+ Cross-encoder re-ranking+ Score normalization + thresholding
ContextRaw chunks+ Compression + deduplication+ Summarization + citation building
SecurityNoneNoneRBAC, metadata filtering, audit logs
Multi-tenancyNoneNoneTenant isolation, per-tenant indexes
EvaluationManualLLM-as-judgeAutomated eval pipeline + dashboards
CachingNoneEmbedding cacheEmbedding + response + semantic cache
MonitoringNoneBasic metricsFull observability (traces, logs, alerts)
Latency2-5s2-5s< 2s with caching
Cost$$$$$$
Best forPrototypes, MVPsProduction appsEnterprise systems

flowchart LR
Q["Do you need\nkeyword matching?"]
Q -->|"Yes"| HYBRID["Hybrid Search\n(BM25 + Vector)"]
Q -->|"No"| VEC["Vector Search Only"]
HYBRID --> Q2["Is accuracy\ncritical?"]
VEC --> Q2
Q2 -->|"Yes"| RERANK["Add Re-ranking\n(Cross-Encoder)"]
Q2 -->|"No"| SIMPLE["Basic Top-K"]
RERANK --> Q3["User queries\nvague?"]
SIMPLE --> Q3
Q3 -->|"Yes"| MQ["Add Multi-Query\n(Query Expansion)"]
Q3 -->|"No"| OK["Good enough"]
MQ --> Q4["Need precise\n+ context?"]
OK --> Q4
Q4 -->|"Yes"| PC["Add Parent-Child\nRetrieval"]
Q4 -->|"No"| DONE["✅ Done"]
style Q fill:#f59e0b,color:#fff
style Q2 fill:#f59e0b,color:#fff
style Q3 fill:#f59e0b,color:#fff
style Q4 fill:#f59e0b,color:#fff
style DONE fill:#22c55e,color:#fff
ScenarioBetter Approach
You need the model to learn new facts permanentlyFine-tuning
You have < 20 documents and the LLM context window fits them allJust put them in the prompt
You need the model to follow a specific output formatFine-tuning or structured output
Your documents change every minuteStreaming data pipeline rather than batch RAG
You need 100% factual accuracy with no hallucination riskConstrained extraction (no generation)
Your users need real-time data from APIsFunction calling / tool use (Phase 6)

  • Separate ingestion and query pipelines
  • Asynchronous ingestion with message queue
  • Stateless query servers behind load balancer
  • Embedding cache (Redis) with TTL
  • Response cache for frequent queries
  • Metadata filtering before vector search
  • Re-ranker between retriever and LLM
  • Authentication (JWT / SSO)
  • Authorization (RBAC with metadata filtering)
  • PII masking before indexing
  • Encryption at rest (AES-256)
  • Encryption in transit (TLS 1.3)
  • Rate limiting per user/API key
  • Audit logging for all queries
  • Secrets management (not in code)
  • Query latency < 3s p95
  • Embedding cache hit rate > 60%
  • Batch vector DB writes (100+ per batch)
  • Connection pooling for all databases
  • Streaming responses for LLM output
  • CDN for static assets and cached responses
  • Horizontal scaling for stateless services
  • Latency metrics per service (p50, p95, p99)
  • Cost tracking per query per user
  • Retrieval precision/recall dashboard
  • LLM hallucination rate monitoring
  • Cache hit/miss ratio
  • Error rate alerts
  • Vector DB usage metrics

FeatureKeyword SearchSemantic SearchHybrid Search
MethodBM25 / TF-IDFVector similarityBM25 + Vector fusion
MatchExact wordsMeaningBoth
Handles typosNoYesYes
Handles synonymsNoYesYes
Handles codeGood (exact API names)PoorGood
Latency< 50ms< 100ms< 200ms
Best forCode, exact termsGeneral text, conceptsProduction systems
AspectDevelopmentProduction
Vector DBLocal (Chroma, FAISS)Managed (Qdrant, Pinecone)
LLMGPT-4o-miniGPT-4o / Claude 3
CachingNoneEmbedding + Response + Semantic
MonitoringConsole logsGrafana + Prometheus + Alerts
SecurityNoneAuth + RBAC + Encryption + Audit
ScalingSingle processHorizontal, auto-scaling
Uptime99%99.9%+
Cost< $100/mo$500-$5000/mo
StrategyBest ForChunk SizeOverlap
Fixed sizeSimple documents256-512 tokens10-20%
RecursiveGeneral purpose512 tokens10-20%
SemanticWell-structured textVariableNone needed
Section-basedDocumentationPer sectionNone needed
Function-levelCodePer functionImports as context
Parent-ChildPrecision + contextSmall child (128), large parent (1024)Hierarchical

flowchart LR
subgraph LEVEL1["Level 1: Foundation ✓"]
L1_1["✅ Embeddings & Vector Search"]
L1_2["✅ Vector Databases"]
L1_3["✅ Basic RAG Pipeline"]
end
subgraph LEVEL2["Level 2: Building ✓"]
L2_1["✅ Chunking Strategies"]
L2_2["✅ Document Ingestion"]
L2_3["✅ Query Pipeline"]
end
subgraph LEVEL3["Level 3: Advanced ✓"]
L3_1["✅ Hybrid Search"]
L3_2["✅ Re-ranking"]
L3_3["✅ Multi-Query & Parent-Child"]
end
subgraph LEVEL4["Level 4: Production ✓"]
L4_1["✅ Production Architecture"]
L4_2["✅ Security & Multi-tenancy"]
L4_3["✅ Evaluation & Monitoring"]
L4_4["✅ Scaling & Caching"]
end
subgraph LEVEL5["Level 5: Projects ✓"]
L5_1["✅ Chat with PDF"]
L5_2["✅ Knowledge Assistant"]
L5_3["✅ Documentation Bot"]
L5_4["✅ Code Assistant"]
end
subgraph NEXT["Next: Phase 6"]
N1["🤖 AI Agents"]
N2["🛠️ Tool Use"]
N3["🧠 Agent Memory"]
N4["👥 Multi-Agent Systems"]
end
LEVEL1 --> LEVEL2 --> LEVEL3 --> LEVEL4 --> LEVEL5 --> NEXT
style LEVEL1 fill:#3b82f6,color:#fff
style LEVEL2 fill:#8b5cf6,color:#fff
style LEVEL3 fill:#f59e0b,color:#fff
style LEVEL4 fill:#22c55e,color:#fff
style LEVEL5 fill:#ef4444,color:#fff
style NEXT fill:#6366f1,color:#fff

ProjectKey LessonArchitecture Pattern
ChatPDFDocument ingestion pipeline + Query pipelineTwo-pipeline architecture
Knowledge AssistantAccess control via metadata filteringPre-filter before search
Documentation BotSection-based chunking + version managementCitation system + version tags
Code AssistantAST-based chunking + dependency graphSymbol index + multi-search
flowchart TD
subgraph COST["Relative Cost to Run"]
P1["📄 ChatPDF\n$200-800/mo"]
P2["🏢 Knowledge Assistant\n$1,000-5,000/mo"]
P3["📚 Documentation Bot\n$100-500/mo"]
P4["💻 Code Assistant\n$500-2,000/mo"]
end
subgraph COMPLEXITY["Implementation Complexity"]
C1["⭐ ChatPDF\nMedium"]
C2["⭐⭐ Knowledge Assistant\nHigh"]
C3["⭐ Documentation Bot\nLow-Medium"]
C4["⭐⭐⭐ Code Assistant\nVery High"]
end
subgraph TIME["Time to MVP"]
T1["📄 ChatPDF\n1-2 weeks"]
T2["🏢 Knowledge Assistant\n4-8 weeks"]
T3["📚 Documentation Bot\n1-2 weeks"]
T4["💻 Code Assistant\n6-12 weeks"]
end
P1 -.-> C1 -.-> T1
P2 -.-> C2 -.-> T2
P3 -.-> C3 -.-> T3
P4 -.-> C4 -.-> T4
style P1 fill:#3b82f6,color:#fff
style P2 fill:#8b5cf6,color:#fff
style P3 fill:#22c55e,color:#fff
style P4 fill:#ef4444,color:#fff
style C1 fill:#3b82f6,color:#fff
style C2 fill:#8b5cf6,color:#fff
style C3 fill:#22c55e,color:#fff
style C4 fill:#ef4444,color:#fff
style T1 fill:#3b82f6,color:#fff
style T2 fill:#8b5cf6,color:#fff
style T3 fill:#22c55e,color:#fff
style T4 fill:#ef4444,color:#fff

You’ve mastered retrieval — giving LLMs access to knowledge. Now learn how to give LLMs access to tools, actions, and autonomy:

  1. Agent Architecture — How agents think: perceive → reason → act → observe
  2. Tool Use — How agents call APIs, run code, and interact with the world
  3. Agent Memory — Short-term, long-term, and episodic memory for agents
  4. Planning — How agents break down complex tasks into steps
  5. Multi-Agent Systems — How agents collaborate, delegate, and debate
  6. Agent Safety — Guardrails, human-in-the-loop, and failure handling

ChunkTopicsDocuments
Chunk 1: Embeddings & Vector SearchWhy retrieval, embeddings, vector space, similarity search, vector DBs01-05
Chunk 2: Building Your First RAG PipelineWhat is RAG, chunking, ingestion pipeline, retrievers, complete pipeline06-10
Chunk 3: Advanced Retrieval TechniquesHybrid search, re-ranking, context compression, parent-child, multi-query11-15
Chunk 4: Production RAG SystemsArchitecture, metadata filtering, evaluation, scaling, best practices16-20
Chunk 5: Enterprise RAG ProjectsChatPDF, Knowledge Assistant, Documentation Bot, Code Assistant, Summary21-25

Previous: 24 — GitHub Code Assistant

Next: Phase 6 — AI Agents & Agentic Systems (Coming Soon)

🎉 Congratulations on completing Phase 5: Retrieval Systems & RAG!