Skip to content

Phase 4 Cheat Sheet

Phase 4: Large Language Models — Cheat Sheet

Section titled “Phase 4: Large Language Models — Cheat Sheet”
ConceptSummary
LLMNeural network trained on internet-scale text to predict next token
AutoregressiveGenerates one token at a time, each depending on all previous
TokenizationText → subword tokens → numerical IDs (vocabulary: 50K-200K)
Context WindowMax tokens model can process (4K → 128K → 1M)
Emergent AbilitiesCapabilities that appear at scale (reasoning, in-context learning)
flowchart LR
subgraph BLOCK["One Transformer Decoder Block"]
IN["Input"] --> ATTN["Multi-Head\nSelf-Attention"]
IN --> RESID1["+"]
ATTN --> RESID1
RESID1 --> LN1["LayerNorm"]
LN1 --> FFN["Feed-Forward\nNetwork"]
LN1 --> RESID2["+"]
FFN --> RESID2
RESID2 --> LN2["LayerNorm"]
LN2 --> OUT["Output"]
end
ComponentPurpose
Self-AttentionComputes relationships between all token pairs
QKVQuery (what I need), Key (what I offer), Value (what I carry)
Multi-Head8-96 parallel attention patterns per layer
Positional EncodingSinusoidal/learned — adds order information
FFNTwo-layer MLP (typically 4x hidden dim)
Residual ConnectionsSkip connections enabling 12-100+ layer depth
LayerNormStabilizes activations, enables training
Stage 1: Pretraining ───→ Stage 2: SFT ───────────→ Stage 3: Alignment
(Trillions of tokens) (100K-1M examples) (RLHF or DPO)
│ │ │
▼ ▼ ▼
Base Model Instruction Model Aligned Model
(continues text) (follows prompts) (helpful + safe)
StageDataObjectiveResult
PretrainingRaw internet textNext token predictionBase model
SFTInstruction-response pairsLanguage modeling on responsesInstruction model
RLHFHuman preference rankingsPPO against reward modelAligned model
DPOPreference pairsDirect preference optimizationAligned model
StrategyBehaviorUse Case
GreedyAlways pick most likely tokenFactual answers, code
Beam SearchMaintain top-N sequencesTranslation, summarization
TemperatureScale logits (0=deterministic, >0=random)Creative writing
Top-KSample from K most likely tokensBalanced creativity
Top-PSample from cumulative probability PAdaptive diversity
  • Temperature: 0.1-0.3 (factual), 0.7-0.9 (creative)
  • Top-K: 40-50
  • Top-P: 0.9-0.95
FeatureHow It Works
StreamingServer-Sent Events (SSE) — tokens sent as generated
Function CallingLLM outputs JSON with tool name + arguments
Structured OutputGrammar-based constrained decoding for valid JSON
Hallucination MitigationRAG, temperature control, prompting, validation
ModelYearParamsContextKey Innovation
GPT-12018117M512Transformer decoder for language modeling
GPT-220191.5B1024Zero-shot generalization
GPT-32020175B2048In-context learning, emergent abilities
GPT-3.52022175B4K/16KInstruction tuning + RLHF
GPT-42023~1.8T8K/32K/128KMultimodal, reasoning
GPT-4o2024—128KReal-time audio, vision, text
FamilyOpen?Best For
GPT (OpenAI)NoGeneral purpose, coding, reasoning
Claude (Anthropic)NoLong documents, safety, analysis
Gemini (Google)NoVery long context, multimodal
Llama (Meta)YesSelf-hosting, research
MistralPartialEfficiency, multilingual
DeepSeekYesCost-effective, coding
Qwen (Alibaba)YesCoding, math, Chinese/English