Skip to content

19. Performance & Scaling

Scaling a RAG system is not about making one component faster — it’s about distributing work across replicated, cached, and optimized services that collectively handle millions of queries with sub-second latency.

Performance and scaling transform a working prototype into a production system that serves thousands of concurrent users without breaking.

flowchart TD
subgraph SMALL["Small Scale (1 user)"]
A["💻 Single Script\n1 retriever\n1 LLM call\nNo caching"]
end
subgraph MEDIUM["Medium Scale (1,000 users)"]
B["🖥️ API Server\n+ Cache Layer\n+ Async Workers"]
end
subgraph LARGE["Large Scale (1M+ users)"]
C["🌐 Load Balancer\n→ API Cluster\n→ Cache Cluster\n→ DB Cluster\n→ LLM Gateway\n→ CDN"]
end
SMALL --> MEDIUM --> LARGE
style SMALL fill:#3b82f6,color:#fff
style MEDIUM fill:#f59e0b,color:#fff
style LARGE fill:#22c55e,color:#fff

A basic RAG pipeline on your laptop:

  • One LLM call per query: 1–5 seconds
  • One embedding computation: 200–500ms (every time)
  • One vector DB search: 50–100ms (on small dataset)
  • No caching: every query pays full cost

At 1 user, this is fine. At 1,000 concurrent users:

  • Queue builds up
  • Response times balloon to 10+ seconds
  • API costs skyrocket
  • Users abandon the application
ProblemSolutionImpact
Slow retrievalCaching, indexing, parallel search10–100x faster
High costResponse caching, embedding reuse, batch processing50–90% cost reduction
Concurrent usersLoad balancing, horizontal scaling, connection poolingHandle 1000x more users
LLM bottlenecksFallback models, streaming, request batchingReliable performance
Data growthSharding, partitioning, tiered storageScale to billions of vectors

Imagine a restaurant that suddenly gets 100x more customers.

Without scaling:

  • One chef tries to cook everything
  • Orders pile up
  • Food takes 45 minutes
  • Customers leave

With scaling:

  • Multiple chefs (horizontal scaling) — each handles different orders
  • Pre-prepped ingredients (caching) — sauces and dough prepared in advance
  • Specialized stations (microservices) — grill station, salad station, dessert station
  • Order prioritization (queues) — simple orders served fast, complex ones take longer
  • Menu optimization (cost management) — focus on high-margin, fast-to-prepare dishes

This is exactly how we scale RAG systems.


Caching is the single most effective optimization for RAG systems.

flowchart LR
Q["🔍 User Query"] --> CACHE_CHECK["💾 Cache Lookup"]
CACHE_CHECK -->|"Cache HIT 🎯"| RESP["✅ Instant Response\n(< 10ms)"]
CACHE_CHECK -->|"Cache MISS ❌"| FULL["🔄 Full Pipeline\n(1-5 seconds)"]
FULL --> STORE["💾 Store in Cache"]
STORE --> RESP
style CACHE_CHECK fill:#f59e0b,color:#fff
style FULL fill:#3b82f6,color:#fff
style RESP fill:#22c55e,color:#fff

Cache the embedding vectors for frequently asked questions.

AspectDetail
What it cachesQuery text → embedding vector mapping
Cache keyHash of query text (normalized: lowercase, stripped)
TTL24–48 hours (embeddings don’t change often)
Hit Rate30–60% for applications with common queries
StorageRedis / Memcached
SavingsSaves 200–500ms per cached query

Cache the complete LLM response for identical questions.

AspectDetail
What it caches(Query + context hash) → response mapping
Cache keyHash of query + retrieved chunk IDs
TTL1–24 hours depending on data freshness needs
Hit Rate20–40% for knowledge-base style applications
StorageRedis cluster with persistence
SavingsSaves 1–5 seconds per cached query + LLM cost

Cache the top-K retrieved documents for common queries.

AspectDetail
What it cachesQuery embedding → top document IDs mapping
Cache keyHash of query embedding vector
TTL1–6 hours (documents may change)
Hit Rate40–60% for FAQ-style knowledge bases
StorageRedis / in-memory
SavingsSaves vector DB query time (50–200ms)

Cache static responses at the edge.

AspectDetail
What it cachesResponses for public, non-personalized queries
LocationEdge servers (Cloudflare, Akamai)
TTL5–60 minutes
Best forPublic documentation Q&A, help center articles

flowchart TD
USER["👤 User"] --> LB["⚖️ Load Balancer"]
LB --> API["📡 API Servers (x10)"]
API --> CACHE1["💾 Response Cache\n(Redis Cluster)"]
API --> CACHE2["💾 Embedding Cache\n(Redis Cluster)"]
CACHE1 -->|"Miss"| RET["📡 Retriever Cluster"]
CACHE2 -->|"Miss"| EMBED["🧠 Embedding Service"]
RET --> VDB[("🗄️ Vector DB Cluster\n(Sharded + Replicated)")]
VDB --> RET
RET --> RERANK["📊 Re-ranker Cluster"]
RERANK --> LLMGW["🤖 LLM Gateway\n(Load Balanced)"]
LLMGW --> LLM1["GPT-4"]
LLMGW --> LLM2["Claude"]
LLMGW --> LLM3["GPT-4o-mini (fallback)"]
LLMGW --> CACHE1
style API fill:#3b82f6,color:#fff
style CACHE1 fill:#f59e0b,color:#fff
style CACHE2 fill:#f59e0b,color:#fff
style VDB fill:#8b5cf6,color:#fff
style LLMGW fill:#22c55e,color:#fff

flowchart TD
USERS["👥 Many Users"] --> LB["⚖️ Load Balancer\nRound Robin / Least Connections"]
LB --> API1["📡 API Server 1"]
LB --> API2["📡 API Server 2"]
LB --> API3["📡 API Server 3"]
LB --> API4["📡 API Server N"]
API1 --> CACHE["💾 Shared Cache\n(Redis Cluster)"]
API2 --> CACHE
API3 --> CACHE
API4 --> CACHE
CACHE --> VDB[("🗄️ Vector DB\n(Read Replicas)")]
style LB fill:#f59e0b,color:#fff
style API1 fill:#3b82f6,color:#fff
style API2 fill:#3b82f6,color:#fff
style API3 fill:#3b82f6,color:#fff
style API4 fill:#3b82f6,color:#fff

Strategy: Add more copies of stateless services (API servers, retrievers, re-rankers) behind a load balancer.

Best for: API servers, retriever instances, re-ranker instances

flowchart TD
Q["🔍 Query"] --> ROUTER["🔀 Query Router\nDetermines which shard based on metadata"]
ROUTER --> S1["📁 Shard 1\nTenants A-H"]
ROUTER --> S2["📁 Shard 2\nTenants I-P"]
ROUTER --> S3["📁 Shard 3\nTenants Q-Z"]
S1 --> R1["✅ Results from Shard 1"]
S2 --> R2["✅ Results from Shard 2"]
S3 --> R3["✅ Results from Shard 3"]
R1 --> MERGE["🔗 Merge Results"]
R2 --> MERGE
R3 --> MERGE
style ROUTER fill:#f59e0b,color:#fff
style S1 fill:#3b82f6,color:#fff
style S2 fill:#3b82f6,color:#fff
style S3 fill:#3b82f6,color:#fff

Strategy: Partition data across multiple database instances based on a shard key (e.g., tenant_id, document hash).

Best for: Vector databases that exceed a single node’s capacity

flowchart TD
WRITER["✍️ Write Node\n(Ingestion only)"] --> REPLICA1["📖 Read Replica 1\n(serves queries)"]
WRITER --> REPLICA2["📖 Read Replica 2\n(serves queries)"]
WRITER --> REPLICA3["📖 Read Replica 3\n(serves queries)"]
Q1["🔍 Query"] --> LB["⚖️ Load Balancer"]
Q2["🔍 Query"] --> LB
Q3["🔍 Query"] --> LB
LB --> REPLICA1
LB --> REPLICA2
LB --> REPLICA3
style WRITER fill:#f59e0b,color:#fff
style REPLICA1 fill:#22c55e,color:#fff
style REPLICA2 fill:#22c55e,color:#fff
style REPLICA3 fill:#22c55e,color:#fff

Strategy: Create read replicas of the vector database. All ingestion goes to the write node; all queries are distributed across read replicas.

Best for: Read-heavy workloads (common in RAG — many queries, fewer document updates)


flowchart LR
Q["🔍 Query Receipt\n0ms"] --> AUTH["🔐 Auth + Filter\n~5ms"]
AUTH --> CACHE["💾 Cache Check\n~5ms"]
CACHE --> RET["📡 Retrieve\n~100ms"]
RET --> RERANK["📊 Re-rank\n~200ms"]
RERANK --> PROMPT["📝 Build Prompt\n~5ms"]
PROMPT --> LLM["🤖 LLM Generate\n~500ms - 2s"]
LLM --> RESP["📨 Response\nTotal: ~1-2.5s"]
style AUTH fill:#3b82f6,color:#fff
style RET fill:#f59e0b,color:#fff
style RERANK fill:#8b5cf6,color:#fff
style LLM fill:#22c55e,color:#fff
TechniqueSavingsImplementation
Embedding cache200–500msCache query embeddings in Redis
Response cache1–5sCache full responses for common queries
Streaming300ms–1s to first tokenStream LLM response instead of waiting for complete
Smaller models200–500ms per callUse GPT-4o-mini for simple queries, GPT-4 for complex
Connection pooling50–100msReuse connections to DB and LLM API
Batch retrieval30–50% reductionBatch multiple queries into one DB call
Approximate search50–80% fasterUse HNSW instead of exact search
Pre-compute100% at query timePre-compute embeddings during ingestion

StrategySavingsTrade-off
Response caching50–90% on LLM costsStale responses for dynamic content
Smaller LLM for simple queries40–60% on LLM costsLower quality on complex queries
Batch embedding30–50% on embedding costsHigher initial latency for batch processing
Token budgeting20–40% on LLM costsNeed to balance context vs. quality
Result count reduction10–20% on LLM costsFewer retrieved chunks → potentially lower quality
Model distillationUp to 90%Need to train and maintain distilled model
Per-Query Cost Distribution:
Embedding API
┌──────────────────┐
│ 5-10% │
└──────────────────┘
┌──────────────────┐
Vector DB │ 10-15% │
└──────────────────┘
┌──────────────────────────────────────┐
LLM (GPT-4o) │ 70-80% │
└──────────────────────────────────────┘
┌──────────────────┐
Infrastructure │ 5-10% │
└──────────────────┘

flowchart TD
PROBLEM["⚠️ High Latency / High Cost"] --> ANALYZE["🔍 Analyze Bottleneck"]
ANALYZE --> CHECK1["Is LLM response slow?"]
ANALYZE --> CHECK2["Is retrieval slow?"]
ANALYZE --> CHECK3["Is re-ranking slow?"]
ANALYZE --> CHECK4["Is cost too high?"]
CHECK1 -->|"Yes"| LLMFIX["🔧 Fix:\n- Use streaming\n- Smaller model\n- Response caching\n- Model fallback"]
CHECK2 -->|"Yes"| RETFIX["🔧 Fix:\n- Embedding cache\n- Vector DB index tuning\n- Connection pooling\n- Read replicas"]
CHECK3 -->|"Yes"| RRFIX["🔧 Fix:\n- Skip re-ranking for simple queries\n- Parallel re-ranking\n- Smaller cross-encoder model"]
CHECK4 -->|"Yes"| COSTFIX["🔧 Fix:\n- Response caching\n- Smaller model routing\n- Token budget reduction\n- Batch processing"]
style PROBLEM fill:#ef4444,color:#fff
style ANALYZE fill:#f59e0b,color:#fff
style LLMFIX fill:#22c55e,color:#fff
style RETFIX fill:#22c55e,color:#fff
style RRFIX fill:#22c55e,color:#fff
style COSTFIX fill:#22c55e,color:#fff

AspectImplementation
ScalingHorizontal scaling with Kubernetes. Multiple retriever instances behind load balancer.
CachingAggressive response caching for common queries (news, trending topics). Embedding cache for repeated queries.
LLM StrategyRoutes simple queries to faster models, complex queries to stronger models. Streaming by default.
AspectImplementation
ScalingAzure infrastructure with global distribution. Data replicated across regions for low latency.
CachingMulti-layer caching: edge cache (CDN) → application cache → database cache.
LLM StrategyOptimized models with batching and prompt caching for efficiency.
AspectImplementation
ScalingAWS-managed auto-scaling. Vector database (Aurora or Pinecone) with read replicas.
CachingBedrock’s managed caching layer. Response caching at the API Gateway level.
LLM StrategyMulti-model support with automatic fallback. Provisioned throughput for predictable workloads.

MistakeWhy It’s WrongFix
Premature optimizationOptimizing before measuring actual bottlenecksProfile first, optimize the actual bottleneck
Ignoring cold startCache hit rates are low initially, causing slow startupPre-warm caches with historical queries
No connection poolingEvery request opens a new connection → TCP overheadUse persistent connection pools
Over-relying on cachingCached responses may be stale for dynamic dataUse appropriate TTLs and include data freshness metadata
Single AZ deploymentEntire system goes down if the availability zone failsDeploy across multiple AZs/regions
No rate limitingOne aggressive user can overwhelm the systemImplement per-user and global rate limits

  1. Measure first, optimize second — Profile your system to find the actual bottleneck before investing in optimizations.

  2. Cache aggressively, invalidate carefully — Caching is the highest-impact optimization. But ensure cache invalidation is correct for your use case.

  3. Design for failure — Every component should have a fallback. If GPT-4 is down, fall back to Claude. If the vector DB is slow, fall back to keyword search.

  4. Scale stateless services easily — API servers, retrievers, and re-rankers should be stateless so they can be scaled horizontally without complexity.

  5. Use connection pooling — Opening new connections to databases and LLM APIs is expensive. Always use connection pools.

  6. Implement graceful degradation — When the system is under load, prioritize quality for paying users and provide degraded service (longer timeout, smaller model) for free tier.

  7. Monitor and auto-scale — Set up auto-scaling based on queue depth, CPU, and request latency. Scale up before problems occur.


Q: Why is caching important for RAG systems?

Caching stores the results of previous computations so they can be reused. In RAG systems, caching can save embedding computations (200–500ms) and LLM calls (1–5 seconds). For common queries, caching can reduce latency by 10–100x and reduce costs by 50–90%. The three main cache types are embedding cache, response cache, and retriever cache.

Q: What is the difference between horizontal and vertical scaling?

Horizontal scaling means adding more machines (e.g., 10 servers instead of 1). Vertical scaling means making a single machine more powerful (e.g., more RAM, faster CPU, better GPU). Horizontal scaling is preferred for RAG systems because it provides fault tolerance, unlimited scalability, and can be automated with load balancers.

Q: How would you design a cache invalidation strategy for a RAG system?

Time-based invalidation: Set TTLs based on data freshness needs — 1 hour for news, 24 hours for documentation, 7 days for reference materials.

Event-based invalidation: When documents are added, updated, or deleted, invalidate related cache entries. Store document IDs with each cached response so you can selectively invalidate.

Version-based invalidation: Include the embedding model version and document index version in the cache key. When either changes, all old cache entries are automatically invalidated.

Lazy invalidation: Don’t actively invalidate. Set short TTLs and let cache misses handle freshness. Simpler but wastes some compute on unnecessary recomputation.

Q: How would you reduce LLM costs in a RAG system without significantly reducing quality?

  1. Response caching — Cache identical questions. Highest impact with no quality loss.
  2. Query classification — Route simple queries (e.g., “What’s your name?”) to cheaper models like GPT-4o-mini, complex queries to GPT-4.
  3. Token budgeting — Limit the number of retrieved chunks sent to the LLM. Top-3 instead of Top-5 if quality doesn’t degrade.
  4. Prompt compression — Use shorter system prompts and instructions. Remove redundant examples.
  5. Batch processing — Combine multiple queries into a single LLM call where possible.

Q: Design a caching strategy for a multi-tenant RAG system that handles both shared and tenant-specific documents.

Cache Architecture:

Layer 1 — Global Cache (shared): Caches responses for queries that are common across all tenants. Cache key: query_hash. Used for public documentation queries. TTL: 1 hour.

Layer 2 — Tenant Cache (isolated): Caches responses specific to each tenant. Cache key: tenant_id + query_hash + document_version. Must never leak between tenants. TTL: 6 hours.

Layer 3 — User Cache (personalized): Caches responses that include user-specific context. Cache key: user_id + tenant_id + query_hash. TTL: 30 minutes.

Invalidation Strategy:

  • Document updates → invalidate all cache entries containing that document
  • Tenant index rebuild → invalidate entire tenant cache
  • User permission change → invalidate user cache

Security: Cache keys must include tenant_id and user_id to prevent cross-tenant cache poisoning. Never store PII in cache values.

Q: Design a globally distributed RAG system that serves users across North America, Europe, and Asia with P95 latency < 1 second.

Architecture:

Regional Deployment: Deploy full stack in 3 AWS regions: us-east-1 (NA), eu-west-1 (EU), ap-southeast-1 (Asia). Each region has independent API servers, cache clusters, and retriever instances.

Vector DB Strategy: Active-active with multi-region replication. Write to local region, async replication to others. Read from local region’s replica. Conflict resolution via timestamp-based last-writer-wins.

Global Load Balancer: Route users to nearest region based on latency (using Route53 latency-based routing).

Caching: Regional cache clusters (Redis in each region). Global cache for shared content (Global Datastore for Redis). Cross-region cache invalidation via SQS.

LLM Strategy: Deploy local LLM endpoints in each region (together.ai, fireworks.ai, or self-hosted). Fallback to API-based models.

Ingestion: Write to primary region (us-east-1), async replicate to all regions. Replication lag: < 5 seconds.

Latency Budget: Auth (5ms) + Cache check (5ms) + Retrieval (100ms) + Re-ranking (200ms) + LLM first token (500ms) = ~810ms P95.

Q: Design a RAG system that can handle 10,000 QPS (queries per second) with a knowledge base of 1 billion documents.

Architecture:

Ingestion Pipeline:

  • Kafka for queueing document processing
  • Spark cluster for parallel chunking + embedding generation
  • Documents partitioned by hash → distributed embedding workers
  • Vectors written to S3 (parquet) → loaded into vector DB via bulk import

Vector Database:

  • 100 nodes, each with 10M vectors (HNSW index)
  • Sharded by document hash across nodes
  • 2x replication for fault tolerance
  • GPU-accelerated search for low latency

Query Pipeline:

  • 500 API servers behind NLB
  • Embedding cache in Redis Cluster (50 nodes)
  • Query router distributes to all 100 vector DB nodes in parallel
  • Each node returns top-10 → merge → re-rank (top-5)
  • LLM gateway with request batching + streaming

Caching:

  • L1: Application-level LRU cache per API server (10K entries, 1 second)
  • L2: Redis cluster (10M entries, 1 hour)
  • L3: CDN for public queries (1M entries, 5 minutes)

Estimated Resources:

  • 500 API servers (c5.xlarge)
  • 100 vector DB nodes (c6i.4xlarge with GPU)
  • 50 Redis nodes (r6g.large)
  • 20 LLM endpoints (A100-80GB)
  • Auto-scaling based on queue depth and CPU

ConceptKey Point
Caching10–100x latency improvement, 50–90% cost reduction
Horizontal scalingAdd more servers behind load balancer
ShardingPartition data across multiple DB instances
ReplicationRead replicas for query scalability
Latency optimizationCache, stream, connection pool, smaller models
Cost optimizationCache, model routing, token budgeting, batch processing
Auto-scalingScale based on demand, not manual intervention
Graceful degradationFall back to cheaper/slower options under load

Previous: 18 — RAG Evaluation & Observability

Next: 20 — Production Best Practices

Related Topics: