Skip to content

Data Modeling

Data modeling is the most critical skill in MongoDB. Good schema design = fast, scalable app. Bad schema = performance nightmares.

3.1 The Core Question: Embed or Reference?

Section titled “3.1 The Core Question: Embed or Reference?”
flowchart TB
subgraph Embed[Embedding ✅
Store related data INSIDE the same document]
E1["user document:
{ name: "Alice",
address: { city: "Mumbai" }
}"]
end
subgraph Reference[Referencing ✅
Store related data in SEPARATE collection, link with ID]
R1["users: { _id: "u1", name: "Alice" }"]
R2["addresses: { userId: "u1", city: "Mumbai" }"]
R1 -.-> R2
end
Q[Which to choose?] --> Embed
Q --> Reference
style Embed fill:#7c3aed,color:#fff
style Reference fill:#3b82f6,color:#fff
Embedding (Denormalization) Referencing (Normalization)
────────────────────────────── ────────────────────────────
Store related data INSIDE Store related data in a
the same document SEPARATE collection and
link with an ID
user document: users collection:
{ { _id: ObjectId("u1"), name: "Alice" }
name: "Alice",
address: { addresses collection:
city: "Mumbai", { userId: ObjectId("u1"),
pin: "400001" city: "Mumbai" }
}
}

Use embedding when:

  • Data is always accessed together
  • Nested data belongs to one parent (ownership)
  • Nested data doesn’t grow unboundedly
  • You need fast reads (one query = complete data)
// ✅ Good embed: User with address (always loaded together, never changes independently)
{
_id: ObjectId("u1"),
name: "Alice Johnson",
email: "alice@example.com",
address: {
street: "123 Marine Drive",
city: "Mumbai",
state: "Maharashtra",
pincode: "400001",
country: "India"
},
preferences: {
theme: "dark",
notifications: true,
language: "en"
}
}
✅ Pros❌ Cons
Single query to load everythingDocument can become very large
Better read performanceUpdating nested data is verbose
Atomic updatesData duplication if shared across documents
Works well for 1-to-1 relationshipsHard to query deeply nested data

Use referencing when:

  • Data is accessed independently
  • Data is shared across many documents
  • Nested data grows unboundedly (e.g., comments on a post)
  • You want to update one place and reflect everywhere
// users collection
{
_id: ObjectId("u1"),
name: "Alice Johnson",
email: "alice@example.com"
}
// orders collection (references user)
{
_id: ObjectId("o1"),
userId: ObjectId("u1"), // reference to user
totalAmount: 2500,
status: "delivered",
createdAt: ISODate("2024-01-10")
}
✅ Pros❌ Cons
No data duplicationRequires multiple queries (or $lookup)
Update once, reflect everywhereSlightly slower reads
Cleaner, smaller documentsNo built-in referential integrity (like SQL)
Works well for growing dataMore complex application code

Example: User and their Profile

One user has exactly one profile.

✅ Recommended: Embed (they’re always loaded together)

// Single document
{
_id: ObjectId("u1"),
email: "alice@example.com",
password: "$2b$10$hashedpassword",
profile: {
firstName: "Alice",
lastName: "Johnson",
bio: "Full-stack developer",
avatar: "https://cdn.example.com/alice.jpg",
dob: ISODate("1995-06-15"),
phone: "+91-9876543210"
},
createdAt: ISODate("2024-01-01")
}

Example: User and their Orders

One user can have many orders.

Option A — Embed orders IN user (❌ Bad for growing data)

// DON'T do this — orders array grows forever, document gets huge
{
_id: ObjectId("u1"),
name: "Alice",
orders: [ ...potentially thousands of orders... ]
// MongoDB document limit = 16MB!
}

Option B — Reference: Store userId in each order (✅ Recommended)

// users collection
{
_id: ObjectId("u1"),
name: "Alice Johnson",
email: "alice@example.com",
totalOrders: 14 // optional: denormalized count for fast access
}
// orders collection
{
_id: ObjectId("o1"),
userId: ObjectId("u1"), // reference back to user
items: [
{ productId: ObjectId("p1"), name: "Laptop Pro", qty: 1, price: 75000 },
{ productId: ObjectId("p2"), name: "Mouse", qty: 2, price: 1200 }
],
totalAmount: 77400,
status: "delivered",
shippingAddress: {
street: "123 Marine Drive",
city: "Mumbai"
},
createdAt: ISODate("2024-01-10")
}

Query: Get all orders for Alice

db.orders.find({ userId: ObjectId("u1") }).sort({ createdAt: -1 })

Example: Students and Courses

One student can enroll in many courses. One course can have many students.

Option A — Reference IDs on both sides:

// students collection
{
_id: ObjectId("s1"),
name: "Ravi Patel",
email: "ravi@example.com",
enrolledCourses: [
ObjectId("c1"), // references to courses
ObjectId("c2"),
ObjectId("c3")
]
}
// courses collection
{
_id: ObjectId("c1"),
title: "Full Stack Web Dev",
instructor: "Prof. Sharma",
price: 4999,
enrolledStudents: [
ObjectId("s1"), // references to students
ObjectId("s2")
]
}

Option B — Junction/Enrollment collection (✅ Better when relationship has extra data)

// students collection
{ _id: ObjectId("s1"), name: "Ravi Patel" }
// courses collection
{ _id: ObjectId("c1"), title: "Full Stack Web Dev" }
// enrollments collection (the "junction table")
{
_id: ObjectId("e1"),
studentId: ObjectId("s1"),
courseId: ObjectId("c1"),
enrolledAt: ISODate("2024-01-05"),
progress: 65, // % completion — belongs to the relationship!
grade: null,
completedAt: null
}

Use a junction collection when the relationship itself has extra data (progress, date enrolled, grade, etc.)


// users
{
_id: ObjectId("u1"),
name: "Alice Johnson",
email: "alice@example.com",
role: "customer",
savedAddresses: [ // embedded — small, user-owned
{
label: "Home",
street: "123 MG Road",
city: "Bangalore",
pincode: "560001"
}
],
createdAt: ISODate("2023-06-01")
}
// products
{
_id: ObjectId("p1"),
name: "MacBook Pro 14",
description: "Apple M3 chip...",
price: 199000,
category: "Laptops",
brand: "Apple",
stock: 25,
images: ["img1.jpg", "img2.jpg"],
tags: ["apple", "laptop", "m3"],
ratings: { average: 4.7, count: 128 }
}
// orders
{
_id: ObjectId("o1"),
userId: ObjectId("u1"), // reference
items: [ // embedded snapshot of product data at time of purchase
{
productId: ObjectId("p1"),
name: "MacBook Pro 14", // snapshot — price might change later!
price: 199000,
quantity: 1
}
],
totalAmount: 199000,
status: "processing", // pending → processing → shipped → delivered
shippingAddress: { // snapshot — user might move later!
street: "123 MG Road",
city: "Bangalore",
pincode: "560001"
},
paymentMethod: "credit_card",
createdAt: ISODate("2024-01-15")
}

💡 Snapshot Pattern: In orders, embed a copy of the product name/price and shipping address. If the product price changes or the user moves, the order history stays accurate.


// users
{
_id: ObjectId("u1"),
name: "Alice",
email: "alice@example.com"
}
// projects
{
_id: ObjectId("proj1"),
name: "Website Redesign",
ownerId: ObjectId("u1"), // reference
members: [ObjectId("u1"), ObjectId("u2"), ObjectId("u3")],
createdAt: ISODate("2024-01-01")
}
// tasks
{
_id: ObjectId("t1"),
title: "Design homepage",
description: "Create wireframes and final design",
projectId: ObjectId("proj1"), // reference
assignedTo: ObjectId("u2"), // reference
createdBy: ObjectId("u1"), // reference
status: "in-progress", // todo → in-progress → review → done
priority: "high",
tags: ["design", "homepage"],
dueDate: ISODate("2024-02-01"),
comments: [ // embedded — always accessed with task
{
userId: ObjectId("u1"),
text: "Please use the brand colors",
createdAt: ISODate("2024-01-15")
}
],
createdAt: ISODate("2024-01-10")
}

flowchart TB
Q1[Is the data always<br/>accessed together?]
Q1 -->|Yes| Q2[Can the nested data<br/>grow without bound?]
Q1 -->|No| Q3[Is the data shared<br/>across many documents?]
Q2 -->|Yes| Q4[Do you need to query<br/>the nested data<br/>independently?]
Q2 -->|No| EMBED["✅ EMBED<br/>Small, bounded, owned data<br/>Single query → fast reads"]
Q3 -->|Yes| REF1["✅ REFERENCE<br/>Shared data → update once<br/>reflected everywhere"]
Q3 -->|No| Q5[Is write performance<br/>critical?]
Q4 -->|Yes| REF2["✅ REFERENCE<br/>Use separate collection<br/>for independent querying"]
Q4 -->|No| Q5
Q5 -->|Yes| EMBED2["✅ EMBED<br/>Fewer documents = less I/O<br/>= faster writes"]
Q5 -->|No| REF3["✅ REFERENCE<br/>Cleaner, normalized data<br/>grows safely"]
style EMBED fill:#7c3aed,color:#fff
style EMBED2 fill:#7c3aed,color:#fff
style REF1 fill:#3b82f6,color:#fff
style REF2 fill:#3b82f6,color:#fff
style REF3 fill:#3b82f6,color:#fff
style Q1 fill:#f59e0b,color:#fff
style Q2 fill:#f59e0b,color:#fff
style Q3 fill:#f59e0b,color:#fff
style Q4 fill:#f59e0b,color:#fff
style Q5 fill:#f59e0b,color:#fff
Ask yourself these questions:
1. Is the data always accessed together?
YES → Consider embedding
NO → Consider referencing
2. Can the nested data grow without bound?
YES (many items over time) → Reference
NO (1 address, 1 profile) → Embed
3. Is the data shared across multiple documents?
YES → Reference (update once)
NO → Embed is fine
4. Do you need to query the nested data independently?
YES → Consider a separate collection
NO → Embedding works
5. Is write performance critical?
YES → Fewer documents (embed) = less I/O
NO → Referencing is fine
Rule of thumb:
─────────────────────────────────────────
Small, bounded, owned data → EMBED
Large, growing, shared data → REFERENCE