Skip to content

02. Embeddings Deep Dive

An embedding is a list of numbers that captures the meaning of a piece of text. Words with similar meanings have similar numbers. That’s how computers understand semantics.

This is one of the most important concepts in AI. Every search engine, recommendation system, and RAG pipeline relies on embeddings. They are the bridge between human language and mathematical computation.


Computers don’t understand words. They understand numbers.

If you give a computer the word “dog”, it sees a string of characters — not the concept of a furry four-legged animal. If you give it “puppy”, it sees another string — with no idea that “puppy” is just a baby dog.

Computers need numbers that carry meaning.

Imagine a huge city. Every restaurant in the city has coordinates — a latitude and longitude.

Thai restaurants tend to be on the same block. Italian restaurants cluster together. Fast food places are across the street from each other.

If you’re standing at a Thai restaurant and want to find another Thai restaurant, you look at nearby coordinates. You don’t walk across town to the Italian district.

Embeddings work exactly like this.

Words and sentences with similar meanings are placed close together in a high-dimensional space. “Dog” and “Puppy” are neighbors. “Dog” and “Car” are far apart.

flowchart TD
subgraph INPUT["Input"]
A["'The cat sat on the mat'"]
B["'A dog played in the park'"]
C["'The stock market rallied today'"]
end
subgraph MODEL["Embedding Model"]
D["🤖 Neural Network\nconverts text → numbers"]
end
subgraph OUTPUT["Output Vectors"]
E["📊 [0.23, 0.87, -0.12, 0.45...]\n(768 numbers)"]
F["📊 [0.21, 0.85, -0.10, 0.42...]\n(768 numbers — similar to E)"]
G["📊 [-0.67, 0.12, 0.89, -0.34...]\n(768 numbers — different from E & F)"]
end
A --> D --> E
B --> D --> F
C --> D --> G
E --> H["'cat' and 'dog' are\nclose together\n(both are pets/animals)"]
G --> I["'stock market' is\nfar away\n(different topic entirely)"]
style INPUT fill:#3b82f6,color:#fff
style MODEL fill:#8b5cf6,color:#fff
style OUTPUT fill:#22c55e,color:#fff

Think of a 2D map, but instead of latitude and longitude, every word has a position based on meaning.

On this map:

  • “Dog” is at position (10, 20)
  • “Puppy” is at position (11, 21) — very close to “Dog”
  • “Cat” is at position (12, 19) — close to both
  • “Car” is at position (80, 90) — far away
  • “Stock” is at position (90, 10) — far from everything

The distance between points represents semantic difference.

In reality, embeddings don’t use 2D. They use 384, 768, 1024, 1536, or even 3072 dimensions. A 2D map captures some meaning but misses nuance. A 1536-dimensional space captures extremely fine-grained semantic relationships.

graph TD
subgraph SEMANTIC["Semantic Space (Simplified)"]
DOG["🐕 Dog"] --- PUPPY["🐶 Puppy"]
DOG --- CAT["🐱 Cat"]
DOG --- ANIMAL["🐾 Animal"]
CAT --- KITTEN["😺 Kitten"]
CAR["🚗 Car"] --- TRUCK["🚛 Truck"]
CAR --- VEHICLE["🚙 Vehicle"]
CAR --- BIKE["🚲 Bike"]
JS["⚡ JavaScript"] --- REACT["⚛️ React"]
JS --- NODE["🟢 Node.js"]
JS --- CODE["💻 Programming"]
end
SEMANTIC --- DISTANCE["✂️ Distance = Semantic Difference"]
style DOG fill:#f59e0b,color:#fff
style CAR fill:#3b82f6,color:#fff
style JS fill:#22c55e,color:#fff

flowchart LR
TEXT["Text Input\n'The capital of France is Paris'"] --> TOKEN["Tokenizer\n(split into tokens)"]
TOKEN --> MODEL["Embedding Model\n(neural network)"]
MODEL --> VECTOR["Embedding Vector\n[0.45, -0.12, 0.78, ...]"]
VECTOR --> USE["Used for:\n• Similarity Search\n• Clustering\n• Classification\n• Retrieval"]
style TEXT fill:#3b82f6,color:#fff
style MODEL fill:#8b5cf6,color:#fff
style VECTOR fill:#22c55e,color:#fff
  1. Input text — Any text: a word, a sentence, a paragraph, or an entire document
  2. Tokenize — Convert text into tokens (subword units)
  3. Embedding model — A neural network processes the tokens and outputs a vector
  4. Vector — A fixed-size array of floating-point numbers representing the text’s meaning

Every dimension in an embedding vector captures a latent feature — something the model learned to track during training. In practice, these features are not human-interpretable. You can’t say “dimension 5 means happiness.” But the pattern of all dimensions together encodes meaning.

DimensionRepresents (hypothetically)High ValueLow Value
1How “animal-like” a word is”dog” = 0.9”table” = 0.1
2How “technology-related""computer” = 0.8”tree” = 0.2
3Sentiment (positive vs negative)“happy” = 0.7”sad” = -0.6
…Hundreds more subtle features……

ModelProviderDimensionsBest ForCost
text-embedding-3-smallOpenAI512-1536General purpose$
text-embedding-3-largeOpenAI256-3072High accuracy$$
voyage-3Voyage AI1024-1536Code + multilingual$$
BGE (BAAI General Embedding)BAAI384-1024Open-sourceFree
E5 (Embedding from Sentences)Microsoft384-768Academic benchmarksFree
all-MiniLM-L6-v2Sentence Transformers384Lightweight, fastFree
Cohere EmbedCohere4096Enterprise$$
GTEAlibaba768-1024MultilingualFree
flowchart TD
subgraph COMPARISON["Model Size vs Quality"]
L1["384d: MiniLM, BGE-small\n🚀 Fastest, 70% accuracy"]
L2["768d: BGE-base, E5\n⚡ Good balance, 80% accuracy"]
L3["1024d: Voyage-2, GTE\n🎯 Strong, 85% accuracy"]
L4["1536d: OpenAI 3-small\n💪 Default choice, 88% accuracy"]
L5["3072d: OpenAI 3-large\n🏆 Best quality, 92% accuracy"]
end
L1 --> L2 --> L3 --> L4 --> L5
style L1 fill:#22c55e,color:#fff
style L3 fill:#f59e0b,color:#fff
style L5 fill:#ef4444,color:#fff

Think of dimensions like resolution on a screen:

  • 384 dimensions — 480p: You see the general shape, but details are blurry
  • 768 dimensions — 720p: Clear enough for most tasks
  • 1536 dimensions — 1080p: Sharp, captures nuanced differences
  • 3072 dimensions — 4K: Maximum detail, but slower and more expensive

Higher dimensions allow the model to distinguish between more subtle differences. “Excited” and “eager” might overlap in 384 dimensions but be separated in 1536 dimensions.

The trade-off: More dimensions = better accuracy but:

  • More storage space
  • Slower search
  • Higher cost (for API-based models)

When you ask Perplexity a question:

  1. Your question is converted to an embedding vector
  2. Perplexity searches millions of web pages for vectors close to yours
  3. The closest matches are retrieved
  4. The LLM reads those pages and generates a response with citations
sequenceDiagram
participant User
participant App as Perplexity App
participant Embed as Embedding API
participant Index as Vector Index
participant LLM as LLM
User->>App: "What is retrieval-augmented generation?"
App->>Embed: Convert question to embedding
Embed-->>App: [0.23, 0.87, -0.12, ...]
App->>Index: Find nearest vectors
Index-->>App: Top 5 relevant web pages
App->>LLM: Question + retrieved pages
LLM-->>App: Answer with citations
App-->>User: "RAG is a technique that... [Source 1] [Source 2]"

MistakeWhy It’s Wrong
❌ “I can use any embedding model interchangeably”Different models have different vector spaces. Vectors from OpenAI cannot be compared directly with vectors from Cohere
❌ “Higher dimensions are always better”Higher dimensions mean more storage and slower search. Choose based on your accuracy vs speed needs
❌ “Embeddings capture exact facts”Embeddings capture meaning, not facts. Two sentences with opposite facts but similar wording will have similar embeddings
❌ “I only need one embedding per document”Long documents need multiple embeddings (one per chunk). A single embedding for a 50-page document loses too much information

Q: What is an embedding in the context of AI?

An embedding is a numerical representation of text — a fixed-size vector of floating-point numbers that captures the semantic meaning. Similar texts have similar embeddings (vectors that are close together in space).

Q: Why do different embedding models produce different vector dimensions? How do you choose?

Different models use different architectures and training data. Higher dimensions (1536-3072) capture more nuance but cost more to store and search. Lower dimensions (384-768) are faster and cheaper but may miss subtle differences. Choose based on your accuracy requirements and latency budget.

Q: You’re building a multilingual retrieval system. How do embeddings handle language differences?

Most embedding models are trained primarily on English. For multilingual search, you need a model specifically trained on multiple languages (e.g., BGE-m3, GTE, Voyage-multilingual). Even then, cross-language search is less accurate than same-language search. A common pattern is to detect the language and route to language-specific embedding models, or use a translation layer before embedding.


ConceptKey Point
EmbeddingA vector of numbers representing text meaning
SimilaritySimilar texts have similar vectors (close together)
DimensionsHigher dimensions = more nuance, more cost
ModelsOpenAI, Voyage, BGE, E5, Cohere — each with different trade-offs
UsageSearch, clustering, classification, RAG

Previous: 01 — Why Retrieval Systems Exist →

Next: 03 — Vector Space →