Module 3: Training
Module 3: Training
Section titled “Module 3: Training”From raw internet text to helpful assistant. Learn the entire training pipeline — pretraining, fine-tuning, RLHF, and alignment.
Overview
Section titled “Overview”Module 3 covers how a raw model becomes a useful assistant. You’ll learn the three-stage training pipeline: pretraining on massive text data, supervised fine-tuning on instruction-response pairs, and alignment with human preferences via RLHF or DPO.
Learning Objectives
Section titled “Learning Objectives”After completing this module, you will be able to:
- ✅ Explain the three stages of LLM training
- ✅ Describe how pretraining works at scale
- ✅ Understand next-token prediction as a training objective
- ✅ Explain supervised fine-tuning (SFT)
- ✅ Describe RLHF and the reward model
- ✅ Understand DPO as a simpler alternative to RLHF
- ✅ Compare training vs inference
Prerequisites
Section titled “Prerequisites”| Requirement | Level |
|---|---|
| Module 2: Transformer Architecture | ✅ Required |
| Understanding of loss functions | ⭐ Recommended |
| Basic ML training concepts | 🔄 Covered in Phase 2 |
Estimated Time
Section titled “Estimated Time”| Activity | Time |
|---|---|
| Reading lessons | 3 hours |
| Practice exercises | 45 minutes |
| Mini quiz | 15 minutes |
| Total | ~4 hours |
Lessons
Section titled “Lessons”| # | Lesson | 🔥 | Description |
|---|---|---|---|
| 13 | Pretraining | 🔥 Must Know | Training on trillions of tokens |
| 14 | Next Token Prediction | 🔥 Must Know | The core training objective |
| 15 | Supervised Fine-Tuning | 🧠 Core Concept | Training on instruction-response pairs |
| 16 | RLHF | 💼 Production | Reinforcement Learning from Human Feedback |
| 17 | DPO | 💼 Production | Direct Preference Optimization |
Training Pipeline
Section titled “Training Pipeline”flowchart LR subgraph STAGE1["Stage 1: Pretraining"] A["Internet Text\n(Trillions of tokens)"] --> B["Objective:\nNext Token Prediction"] B --> C["Base Model\n(Good at continuation,\nnot instruction following)"] end
subgraph STAGE2["Stage 2: Supervised Fine-Tuning"] C --> D["Instruction-Response\nPairs (100K-1M)"] D --> E["SFT Model\n(Can follow\ninstructions)"] end
subgraph STAGE3["Stage 3: Alignment"] E --> F["RLHF or DPO"] F --> G["Aligned Model\n(Helpful, harmless,\nhonest)"] end
style STAGE1 fill:#3b82f6,color:#fff style STAGE2 fill:#8b5cf6,color:#fff style STAGE3 fill:#22c55e,color:#fffKey Concepts
Section titled “Key Concepts”- Pretraining: Self-supervised learning on raw text — no labels needed
- Scaling laws: Model performance improves predictably with more data, parameters, and compute
- SFT: Teaching the model to follow instructions using human-written examples
- Reward Model: A separate model trained to predict human preferences
- RLHF: Using PPO to optimize the LLM against the reward model
- DPO: Directly optimizing preferences without a separate reward model
- Alignment: Making models helpful, harmless, and honest (the “HHH” framework)
Module Summary
Section titled “Module Summary”In this module, you learned:
- Pretraining trains the model on raw text using next-token prediction — this is ~99% of the compute
- SFT teaches instruction following using human-written examples
- RLHF uses a reward model trained on human preferences, then PPO to optimize the LLM
- DPO simplifies alignment by directly optimizing preference probabilities
Practice Questions
Section titled “Practice Questions”- Why is pretraining called “self-supervised”? Where do the labels come from?
- What happens if you skip the alignment stage (SFT + RLHF/DPO)?
- Compare RLHF and DPO — what are the advantages of each?
- Why does SFT use a small dataset (100K-1M examples) compared to pretraining (trillions of tokens)?
Interview Questions
Section titled “Interview Questions”-
Q: Why can’t we just use more SFT data instead of RLHF?
- A: SFT teaches the model to mimic the format, but RLHF teaches it to optimize for human preferences. SFT on bad examples can make the model worse; RLHF handles nuanced trade-offs.
-
Q: What is the “alignment tax”?
- A: Alignment (SFT + RLHF/DPO) can reduce the model’s diversity and creativity slightly. The model becomes safer but may produce less varied outputs. This trade-off is called the alignment tax.
Next Steps
Section titled “Next Steps”➡️ Continue to Module 4: Inference →