Skip to content

13. Context Compression

Context compression is the process of reducing retrieved documents to only their most relevant parts before sending them to the LLM. It saves tokens (reducing cost), improves quality (less noise), and fits more useful information into the context window.

Retrieved documents often contain irrelevant content. A 500-token chunk might have only 50 tokens of useful information. Context compression extracts those 50 tokens and discards the rest — making every token in the prompt count.


You search for “What is the capital of France?” The retriever finds a 500-token paragraph about France that includes history, culture, and geography — with the capital information buried in the middle.

You don’t need the whole paragraph. You need: “The capital of France is Paris.”

Context compression extracts just that sentence. Less noise for the LLM. Lower cost. Better answer.


Imagine you’re a CEO who needs to review 20 reports. You don’t have time to read all of them. You ask your assistant to:

  1. Extract — pull the key findings from each report
  2. Summarize — condense each one to 2-3 sentences
  3. Deduplicate — remove information that appears in multiple reports
  4. Prioritize — order by importance

That’s context compression. The LLM (CEO) gets only what it needs, in the right order.

flowchart LR
subgraph BEFORE["Before Compression"]
CHUNKS["5 Chunks:\n1,200 total tokens\nMixed relevance\nContains noise"]
end
subgraph COMPRESSION["Compression Pipeline"]
EXTRACT["📋 Extract\n(key sentences)"]
SUMMARIZE["📝 Summarize\n(condense)"]
DEDUP["🔁 Deduplicate\n(remove repeats)"]
FILTER["🗑️ Filter\n(remove irrelevant)"]
end
subgraph AFTER["After Compression"]
OUTPUT["Compressed Context:\n300 tokens\nHighly relevant\nClean signal"]
end
CHUNKS --> EXTRACT --> SUMMARIZE --> DEDUP --> FILTER --> OUTPUT
style BEFORE fill:#ef4444,color:#fff
style COMPRESSION fill:#8b5cf6,color:#fff
style AFTER fill:#22c55e,color:#fff

Select the most important sentences from each chunk without modifying them.

Input: "France is a country in Western Europe. The capital of France is Paris.
It is known for its cuisine, art, and fashion. The Eiffel Tower is
located in Paris. France has a population of approximately 67 million."
Output: "The capital of France is Paris. The Eiffel Tower is located in Paris."

Pros: Preserves original wording, no hallucination risk Cons: May lose nuance, sentences need to be self-contained

Use an LLM to rewrite the retrieved content into a concise summary.

Input: [5 chunks about React hooks]
Output: "React hooks (useState, useEffect, useContext) allow functional
components to manage state and side effects. Introduced in React 16.8."

Pros: Dense, well-structured, removes all redundancy Cons: LLM cost, potential hallucination in the summary

Remove chunks or sentences that fall below a relevance threshold.

flowchart TD
CHUNKS["Retrieved Chunks:\nChunk A: score 0.92\nChunk B: score 0.87\nChunk C: score 0.45\nChunk D: score 0.34\nChunk E: score 0.12"] --> FILTER2["Filter:\nRemove chunks\nbelow 0.70"]
FILTER2 --> KEPT["Kept:\nChunk A: 0.92\nChunk B: 0.87"]
FILTER2 --> DISCARD["Discarded:\nChunk C: 0.45\nChunk D: 0.34\nChunk E: 0.12"]
style CHUNKS fill:#3b82f6,color:#fff
style KEPT fill:#22c55e,color:#fff
style DISCARD fill:#ef4444,color:#fff

Ask the LLM to “only use the relevant parts of the context” in its prompt.

System: "Answer the question using ONLY the information in the context below.
Ignore any irrelevant information."
Context: [5 chunks with mixed relevance]

The LLM naturally ignores irrelevant parts — but still pays for processing them.


flowchart LR
subgraph IMPACT["Compression Impact"]
NOCOMP["❌ No Compression\n5 chunks × 500 tokens\n= 2,500 input tokens\n💰 $0.00375 per query\n⚠️ More noise for LLM"]
COMPRESSED["✅ With Compression\n~600 compressed tokens\n💰 $0.0009 per query\n🟢 Clean signal for LLM"]
end
NOCOMP --> COMPARISON["60% cost reduction\nImproved answer quality"]
COMPRESSED --> COMPARISON
style NOCOMP fill:#ef4444,color:#fff
style COMPRESSED fill:#22c55e,color:#fff
MetricWithout CompressionWith CompressionImprovement
Input tokens2,50060076% reduction
Cost per query$0.00375$0.0009076% cheaper
LLM accuracy85%92%+7%
Hallucination rate8%4%-50%

One common question: “Why compress when models now support 128K-200K context windows?”

flowchart LR
Q["Do I need\nCompression?"] --> CHECK["What's my\naverage retrieval\ntoken count?"]
CHECK -->|< 4K tokens| NOCOMP["Probably not\n(128K context is enough)"]
CHECK -->|> 4K tokens| YESCOMP["Yes!\nCost + Quality benefits"]
style Q fill:#3b82f6,color:#fff
style NOCOMP fill:#f59e0b,color:#fff
style YESCOMP fill:#22c55e,color:#fff
AspectLong ContextCompression
Context window128K-200K tokensReduces to 500-2000 tokens
CostHigh (more tokens = more $)Low (fewer tokens = less $)
Quality”Lost in the middle” problemOnly relevant info reaches LLM
LatencySlower (more tokens to process)Faster
Best forWhen ALL context is neededWhen only key info is needed

ProductCompression Strategy
PerplexityExtractive — highlights key sentences from sources
Claude ProjectsLLM-based — Claude naturally focuses on relevant context
Notion AISummarization — condenses retrieved notes before answering
Microsoft CopilotFilter-based — removes low-relevance results before generation

PracticeWhy
Compress earlyReduce token count before sending to the LLM to save cost and reduce latency
Combine techniquesUse extractive + filter-based for speed; add LLM-based for maximum compression
Track compression ratioMonitor how many tokens are being removed. 50-70% compression is typical
Don’t over-compressRemoving too much context can lose important nuance. Benchmark with your data

MistakeWhy It’s Wrong
❌ “I don’t need compression with 128K context”Larger context doesn’t mean better results. LLMs perform worse with very long contexts (the “lost in the middle” phenomenon)
❌ “I’ll compress everything with an LLM”LLM-based compression adds latency and cost. Use it sparingly — extractive methods are cheaper and faster for most cases
❌ “Compression might lose important information”Test compression with your data. When tuned properly, compression improves quality by removing noise and focusing the LLM

Q: What is context compression in a RAG system?

Context compression reduces the amount of text sent to the LLM by extracting only the relevant parts, summarizing, filtering out low-quality chunks, and removing duplicate information. It saves tokens (cost) and improves accuracy (less noise).

Q: Compare extractive and abstractive compression. When would you use each?

Extractive compression selects important sentences from the original text — fast, preserves facts, no hallucination risk. Use for factual Q&A where precision matters. Abstractive compression uses an LLM to rewrite/condense the content — more compact but slower and risks hallucination. Use for summarization tasks where brevity matters more than exact wording.

Q: Design a compression strategy for a RAG system that processes 10K queries/day with a $500/month LLM budget.

Strategy: (1) Retrieve — top 10 chunks. (2) Filter — remove chunks below 0.7 similarity threshold (removes ~30% immediately). (3) Deduplicate — remove chunks with >90% semantic overlap. (4) Extractive compression — use sentence-level relevance scoring (e.g., BERT-score) to keep only the top 50% of sentences. (5) Token budget — cap compressed output at 2000 tokens. (6) Monitor — track tokens saved vs answer quality. This should reduce token usage by ~60% while maintaining or improving quality.


ConceptKey Point
Context CompressionReducing retrieved content to essential parts
ExtractiveSelect important sentences (fast, faithful)
AbstractiveLLM rewrites content (dense, risk of hallucination)
Filter-basedRemove low-relevance chunks entirely
Why it mattersLower cost, better quality, faster responses

Previous: 12 — Re-ranking →

Next: 14 — Parent-Child Retrieval →