Skip to content

Module 4: Inference

Everything that happens when you send a prompt to an LLM — from decoding strategies to streaming, function calling, and managing hallucinations.


Module 4 is about the inference pipeline — what happens after training is complete and you send a prompt. You’ll learn how text is generated token by token, how different decoding strategies affect output quality, how streaming works, how function calling enables tool use, and why hallucinations happen.


After completing this module, you will be able to:

  • ✅ Describe the inference pipeline step by step
  • ✅ Compare greedy decoding, beam search, and sampling
  • ✅ Explain how temperature, top-k, and top-p affect output
  • ✅ Understand how streaming delivers tokens in real-time
  • ✅ Explain function calling architecture
  • ✅ Describe structured output methods (JSON mode, grammar-based)
  • ✅ Understand why hallucinations occur and mitigation strategies

RequirementLevel
Module 3: Training✅ Required
Understanding of probability⭐ Recommended
API usage experience🔄 Helpful

ActivityTime
Reading lessons3 hours
Practice exercises45 minutes
Mini quiz15 minutes
Total~4 hours

#Lesson🔥Description
18Inference🔥 Must KnowThe complete inference pipeline
19Decoding Strategies🧠 Core ConceptGreedy, beam search, top-k, top-p
20Temperature, Top-K & Top-P🧠 Core ConceptControlling randomness in generation
21Streaming💼 ProductionReal-time token delivery
22Function Calling💼 ProductionConnecting LLMs to external tools
23Structured Output💼 ProductionJSON mode, constrained generation
24Hallucinations🧠 Core ConceptWhy models fabricate and how to reduce it

flowchart TD
IN["User Prompt"] --> TOK["Tokenizer\n(Text → Token IDs)"]
TOK --> TF["Transformer\n(Multiple decoder blocks)"]
TF --> LOGITS["Logits\n(Unnormalized scores)"]
LOGITS --> SAMPLING["Decoding Strategy\n(Greedy / Sampling / Beam)"]
SAMPLING --> NEXT["Next Token ID"]
NEXT --> DETOK["Detokenizer\n(Token ID → Text)"]
DETOK --> OUT["Output Text"]
NEXT --> APPEND["Append to Input"]
APPEND --> TOK
OUT --> STREAM{"Streaming?"}
STREAM -->|"Yes"| EMIT["Emit token to client"]
STREAM -->|"No"| WAIT["Wait for complete response"]
style IN fill:#3b82f6,color:#fff
style TOK fill:#8b5cf6,color:#fff
style TF fill:#f59e0b,color:#fff
style LOGITS fill:#ef4444,color:#fff
style SAMPLING fill:#22c55e,color:#fff
style OUT fill:#3b82f6,color:#fff

  • Autoregressive generation: Each token depends on all previous tokens
  • Decoding strategies: Greedy (deterministic), beam search (multiple paths), sampling (stochastic)
  • Temperature: Controls the “sharpness” of the probability distribution
  • Top-k: Only sample from the k most likely tokens
  • Top-p (nucleus): Only sample from tokens whose cumulative probability exceeds p
  • Streaming: Server-Sent Events (SSE) deliver tokens as they’re generated
  • Function calling: LLM outputs structured arguments for predefined functions
  • Constrained decoding: Grammar-based generation that guarantees valid output format
  • Hallucinations: Model produces plausible-sounding but incorrect information

In this module, you learned:

  1. Inference is autoregressive — one token at a time, each depending on previous tokens
  2. Decoding strategies balance creativity vs coherence
  3. Temperature, top-k, top-p control output randomness
  4. Streaming improves user experience with real-time output
  5. Function calling enables LLMs to interact with external systems
  6. Structured output ensures machine-parseable responses
  7. Hallucinations are inherent to LLMs and require systematic mitigation

  1. Why is greedy decoding not always the best strategy for creative tasks?
  2. Explain the difference: what happens when temperature = 0 vs temperature = 1 vs temperature = 2?
  3. How would you design a function calling schema for a weather API?
  4. What are three strategies to reduce hallucinations?

  1. Q: Why can’t LLMs generate text in parallel? Why is it sequential?

    • A: Because each token depends on all previous tokens. You can’t know token #100 without first generating tokens 1-99. This is the fundamental constraint of autoregressive generation.
  2. Q: How does temperature 0 differ from greedy decoding?

    • A: At temperature → 0, the softmax approaches argmax, which is equivalent to greedy decoding. Both select the most likely token deterministically.
  3. Q: What are the security implications of function calling?

    • A: Unrestricted function calling can lead to prompt injection where users trick the model into calling sensitive functions. Mitigations include input sanitization, least-privilege function design, and human-in-the-loop for destructive operations.

➡️ Continue to Module 5: Revision & Project →