Skip to content

01. What is a Large Language Model?

A Large Language Model (LLM) is a neural network trained on massive amounts of text to predict the next word in a sequence — and that simple ability, at massive scale, produces systems that can write, reason, translate, summarize, and hold conversations.

LLMs are the technology behind ChatGPT, Claude, Gemini, GitHub Copilot, and every modern AI assistant you have heard of. They represent the culmination of decades of AI research, combining the scale of the internet’s text data with the power of Transformer neural networks and the parallel processing of modern GPU hardware.

graph TD
AI["🤖 Artificial Intelligence\n(Broad field: machines acting smart)"]
ML["📊 Machine Learning\n(Learn from data without explicit rules)"]
DL["🧠 Deep Learning\n(Learn via multi-layer neural networks)"]
TF["⚡ Transformer\n(Attention-based architecture from 2017)"]
LLM["💬 Large Language Model\n(Transformer trained on internet-scale text)"]
AI --> ML
ML --> DL
DL --> TF
TF --> LLM
style AI fill:#3b82f6,color:#fff
style ML fill:#8b5cf6,color:#fff
style TF fill:#f59e0b,color:#fff
style LLM fill:#22c55e,color:#fff

Before LLMs, computers could not understand natural language. They could process text — search for keywords, match patterns, follow rules — but they could not understand meaning.

Traditional NLP (Pre-2018):

  • Keyword matching: “I’m not happy” → searches for “happy” → misses the “not”
  • Rule-based systems: Thousands of hand-written grammar rules that break on every edge case
  • Statistical models: Naive Bayes, SVM — required manual feature engineering, could only handle simple tasks
  • No concept of context, sarcasm, analogy, or reasoning

Example — Customer support before LLMs:

User: “My order hasn’t arrived and it’s been three weeks. I’m extremely frustrated.”

Traditional bot: “I found the keyword ‘order’. Did you mean to track your order? Type your order number.”

The bot could not detect frustration, understand the time reference, or empathize.

LLMs brought three breakthroughs:

  1. Understanding context — They don’t match keywords; they understand the meaning behind words
  2. Generalization — One model can summarize, translate, code, write poetry, and answer questions — no separate model needed for each task
  3. Emergent abilities — At sufficient scale, models develop abilities that were never explicitly programmed: reasoning, analogies, chain-of-thought, even theory of mind
flowchart LR
subgraph BEFORE["Before LLMs (Pre-2018)"]
A1["Task 1: Sentiment\n→ separate model"]
A2["Task 2: Translation\n→ separate model"]
A3["Task 3: Summarization\n→ separate model"]
A4["Task 4: Q&A\n→ separate model"]
end
subgraph AFTER["After LLMs (2020+)"]
B["One Large Language Model\n(pretrained on internet text)"]
B --> C1["Sentiment"]
B --> C2["Translation"]
B --> C3["Summarization"]
B --> C4["Q&A"]
B --> C5["Code"]
B --> C6["Chat"]
end
style BEFORE fill:#ef4444,color:#fff
style AFTER fill:#22c55e,color:#fff
style B fill:#8b5cf6,color:#fff

Every word in the name tells you something important.

WordMeaningWhy It Matters
LargeThe model has billions of parameters and was trained on massive dataSmall models can’t do what large models can — scale creates emergent abilities
LanguageThe model processes human language — text in, text outIt understands words, sentences, documents, and conversations
ModelA mathematical approximation of the patterns in its training dataIt’s not a database, it’s a pattern predictor — it models the probability distribution of language

Put together: A Large Language Model is a mathematical system, trained on billions of text documents, that predicts the next word in a sequence with such accuracy that it appears to understand and generate human language.

flowchart LR
A["LARGE\n(billions of parameters)"] --> D["🧠"]
B["LANGUAGE\n(trained on internet text)"] --> D
C["MODEL\n(probability distribution)"] --> D
D --> E["Ability to generate\nhuman-quality text"]
D --> F["Ability to understand\ncontext and meaning"]
D --> G["Ability to reason\nand follow instructions"]
style A fill:#ef4444,color:#fff
style B fill:#3b82f6,color:#fff
style C fill:#f59e0b,color:#fff
style D fill:#8b5cf6,color:#fff
style E fill:#22c55e,color:#fff
style F fill:#22c55e,color:#fff
style G fill:#22c55e,color:#fff

Imagine a librarian who has read every book ever written — every novel, textbook, blog post, poem, research paper, Wikipedia article, and Reddit thread.

You give her the start of a sentence: “The cat sat on the…”

She doesn’t think about cats or mats. She has read “The cat sat on the mat” 10,000 times, “The cat sat on the roof” 500 times, and “The cat sat on the fence” 200 times.

Based on everything she has read, she predicts the most likely next word: “mat.”

Now expand this: She doesn’t just predict one word. She predicts the next word after that. And the next. One word at a time, she builds entire paragraphs, essays, stories — all based on the statistical patterns she absorbed from her training.

The key insight: She is not thinking. She is not conscious. She is doing what she has done a trillion times: pattern matching at an unimaginable scale.


The Difference Between NLMs, ML, DL, Transformers, and LLMs

Section titled “The Difference Between NLMs, ML, DL, Transformers, and LLMs”
TechnologyWhat It DoesExampleLimitation
Traditional NLPFollows hand-written rules for textRegex, keyword matchingCannot handle ambiguity or nuance
Machine LearningLearns patterns from labeled dataSpam classifier, regressionRequires feature engineering; limited on text
Deep LearningMulti-layer neural networks learn hierarchical featuresCNNs for images, RNNs for sequencesRNNs are slow; can’t parallelize
TransformerArchitecture using self-attention to process sequences in parallelBERT (understanding), GPT (generation)The foundation, not the final product
LLMA massive Transformer trained on internet-scale textChatGPT, Claude, GeminiExpensive to run; can hallucinate
gantt
title Evolution of Language AI
dateFormat YYYY
axisFormat %Y
section Traditional
Rule-based NLP :1950, 50y
Statistical NLP :1990, 20y
ML-based NLP :2000, 15y
section Deep Learning
RNNs / LSTMs :2013, 5y
Attention Mechanism :2015, 2y
section Transformers
Transformer Paper :2017, 1y
BERT / GPT-1 :2018, 1y
GPT-2 :2019, 1y
GPT-3 :2020, 1y
ChatGPT / GPT-4 :2022, 2y
Claude / Gemini / Llama :2023, 2y

flowchart TD
IN["Input Text\n(User prompt or query)"] --> TK["Tokenizer\n(Text → Numbers)"]
TK --> TF["Transformer\n(Billions of parameters)"]
TF --> PR["Prediction\n(Probability over all tokens)"]
PR --> TK2["Sample next token\n(Choose the next word)"]
TK2 --> D{"Is response\ncomplete?"}
D -->|"No — repeat"| IN
D -->|"Yes"| OUT["Final Response\n(Generated text)"]
style IN fill:#3b82f6,color:#fff
style TK fill:#8b5cf6,color:#fff
style TF fill:#f59e0b,color:#fff
style PR fill:#ef4444,color:#fff
style TK2 fill:#8b5cf6,color:#fff
style OUT fill:#22c55e,color:#fff
  1. Input — A user types a prompt: “Explain quantum computing in simple terms”
  2. Tokenization — The text is split into tokens (words/subwords) and converted to numbers
  3. Transformer — The numbers pass through 100+ layers of neural network, each layer refining the representation using self-attention and feed-forward computation
  4. Prediction — The final layer outputs a probability distribution over the entire vocabulary (~50,000–200,000 possible tokens)
  5. Sampling — The model selects one token (the most likely, or randomly weighted by probability)
  6. Repeat — The new token is appended to the input, and steps 2–5 repeat until the model outputs a stop token or reaches the token limit

Here is a comparison of the major LLM families available today:

mindmap
root((LLM Families))
OpenAI
GPT-3.5
GPT-4
GPT-4o
o1 / o3 (reasoning)
Best overall quality
Anthropic
Claude 3 Haiku
Claude 3 Sonnet
Claude 3 Opus
Claude 3.5 / 4
Focus on safety
Google DeepMind
Gemini 1.5 Pro
Gemini 1.5 Flash
Gemini 2.0
1M token context
Meta
LLaMA 2
LLaMA 3 / 3.1
Open-weight
Good for self-hosting
Mistral AI
Mistral 7B
Mixtral 8x7B
Mistral Large
Efficient architectures
Alibaba
Qwen / Qwen 2
Strong on coding
Multilingual
DeepSeek
DeepSeek V2 / V3
R1 (reasoning)
Extremely cost-efficient
Microsoft
Phi 3 / 3.5
Small but capable
Runs on phones
ModelCreatorSize RangeContext WindowCostBest ForOpen?
GPT-4oOpenAIUnknown128K$$$General, coding, reasoningNo
Claude 3 OpusAnthropicUnknown200K$$$Long docs, analysis, safetyNo
Gemini 1.5 ProGoogleUnknown1M$$Very long context, multimodalNo
LLaMA 3.1 405BMeta8B, 70B, 405B128KFreeSelf-hosting, researchYes
Mistral LargeMistralUnknown128K$$Efficiency, multilingualNo
Mixtral 8x7BMistral46B total (MoE)32KLowSelf-hosting on modest hardwareYes
Qwen 2.5 72BAlibaba7B, 14B, 72B128KLow-MedCoding, math, Chinese/EnglishYes
DeepSeek V3DeepSeek671B (MoE)128KVery LowCost-efficient, codingYes
Phi 3Microsoft3.8B, 14B128KVery LowOn-device, simple tasksYes
Gemma 2Google2B, 9B, 27B8KFreeLightweight researchYes

Note on “Open”: Strictly open-source models (open weights + open data + open training code) are rare. Most models listed as “Yes” here are “open-weight” — you can download and run the model weights, but the training data and training code are not public.


flowchart LR
A["Input: 'What is the\ncapital of France?'"] --> B["LLM doesn't 'know'\nParis is the capital"]
B --> C["LLM has seen\n'capital of France → Paris'\nmillions of times"]
C --> D["LLM predicts 'Paris'\nwith 99.9% probability"]
D --> E["You: 'Wow, it knows\ngeography!'"]
E --> F["Reality: It's just\na statistical pattern"]
style A fill:#3b82f6,color:#fff
style D fill:#22c55e,color:#fff
style F fill:#ef4444,color:#fff

LLMs feel intelligent because:

  1. Pattern completion at scale — When you see a friend and say “Long time, no…” you complete “see” automatically. LLMs do this with every concept, across billions of patterns, simultaneously.

  2. Emergent abilities — When a model reaches a certain size (~70B+ parameters), it develops abilities that smaller versions don’t have: reasoning, step-by-step thinking, analogical reasoning. No one programmed these — they emerged from scale.

  3. Massive training data — The model has been exposed to more text than any human could read in 100 lifetimes. It has seen every argument, every explanation, every analogy — and can recombine them.

  4. In-context learning — LLMs can learn a new task from a few examples in the prompt, without any weight updates. This makes them appear to “understand” instructions instantly.

The Critical Distinction: Prediction vs. Understanding

Section titled “The Critical Distinction: Prediction vs. Understanding”

LLMs are prediction engines — not thinking machines.

  • They do not understand truth
  • They do not have beliefs
  • They do not have consciousness
  • They do not have intentions
  • They do not have memory (beyond the context window)

When an LLM says “I think the answer is X,” it is not thinking. It is generating a sequence of tokens that statistically matches the pattern of a thoughtful person giving an answer.


CapabilityExampleWhy It Works
Summarization”Summarize this 100-page PDF”Training data contained countless examples of summaries
Translation”Translate to Spanish”Training data was multilingual
Code generation”Write a Python function to sort a list”Training data contained GitHub repositories
Creative writing”Write a poem about AI in the style of Shakespeare”Training data contained literature in every style
ReasoningSolve a logic puzzle step by stepAt scale, models learn to decompose problems
Role-playing”Act as a history professor”Training data contained dialogues, role-play, and instruction examples
Tool use”Find today’s weather” (calls a function)Fine-tuned for function calling
LimitationExampleWhy
HallucinationStates fake facts confidentlyModel prioritizes plausible text over truth
No true reasoningFails on simple math if pattern is unusualIt’s pattern matching, not mathematical reasoning
No memory beyond contextForgets what you said 1000 tokens agoFixed context window size
No real-time awarenessDoesn’t know today’s news (unless using tools)Training data has a cutoff date
Biases from training dataExhibits social biasesLearns biases present in internet text
No causal understandingCan’t tell correlation vs causationPredicts text, doesn’t model the world
Sensitive to prompt wordingSame question rephrased → different answerNo grounded understanding of meaning

# The simplest way to use an LLM — through an API
# Install: pip install openai
from openai import OpenAI
client = OpenAI(api_key="your-api-key") # Get key from platform.openai.com
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are a helpful tutor who explains things simply."},
{"role": "user", "content": "What is an LLM in one paragraph?"}
]
)
print(response.choices[0].message.content)
# LLM stands for Large Language Model — a neural network trained on massive text data
# that predicts the next word in a sequence. By repeating this prediction millions of
# times, LLMs can generate coherent text, answer questions, write code, and hold
# conversations. They don't "think" or "understand" — they excel at statistical
# pattern matching at an enormous scale.
// Install: npm install openai
import OpenAI from 'openai';
const openai = new OpenAI({ apiKey: 'your-api-key' });
async function askLLM() {
const response = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [
{ role: 'system', content: 'You explain concepts simply.' },
{ role: 'user', content: 'What is an LLM in one paragraph?' }
]
});
console.log(response.choices[0].message.content);
// LLM stands for Large Language Model...
}
askLLM();

flowchart TD
subgraph CLIENT["Client Layer"]
UI["User Interface\n(Chat, API call, IDE plugin)"]
end
subgraph API["API / Gateway Layer"]
AUTH["Authentication\n& Rate Limiting"]
RTR["Router\n(selects model endpoint)"]
end
subgraph INFERENCE["Inference Layer"]
TOK["Tokenizer\n(text → token IDs)"]
TF["Transformer\n(autoregressive generation)"]
SAM["Sampling\n(top-k, top-p, temperature)"]
DETOK["Detokenizer\n(token IDs → text)"]
end
subgraph MODEL["Model Storage"]
W["Model Weights\n(stored on GPU VRAM or disk)"]
CFG["Config\n(architecture settings)"]
end
UI --> AUTH
AUTH --> RTR
RTR --> TOK
TOK --> TF
W --> TF
CFG --> TF
TF --> SAM
SAM --> DETOK
DETOK --> UI
style CLIENT fill:#3b82f6,color:#fff
style API fill:#8b5cf6,color:#fff
style INFERENCE fill:#f59e0b,color:#fff
style MODEL fill:#22c55e,color:#fff

  1. Choose the right model for the task — GPT-4o is great for complex reasoning; smaller models like Claude Haiku or GPT-4o mini are faster and cheaper for simple tasks
  2. Match model size to hardware — Don’t try to run a 405B parameter model on a laptop; use API-based models or quantized smaller models for local use
  3. Understand the pricing model — LLM APIs charge per token (input + output); cost can accumulate fast with long prompts or high traffic
  4. Use system prompts — Set the behavior, tone, and constraints in a system message rather than relying on the user to describe what they want
  5. Always validate outputs — LLMs can hallucinate; verify facts, test code, review generated content before using it
  6. Monitor for prompt injection — Users may try to override your system instructions; implement input sanitization and output filtering

MisconceptionTruth
”LLMs understand language like humans”LLMs have no understanding, consciousness, or awareness — they perform statistical pattern matching
”LLMs have access to the internet”No — they only know what was in their training data (cutoff date varies)
“LLMs will replace all jobs”LLMs are tools that augment humans, not replace them — they lack agency, physical presence, and true reasoning
”Bigger models are always better”Larger models are more capable but also more expensive and slower; smaller models can outperform on specific tasks with fine-tuning
”Open-source LLMs are as good as GPT-4”Open models are catching up fast but typically lag 6–18 months behind the best proprietary models
”LLMs can do math perfectly”LLMs are not calculators — they can approximate but may fail on simple arithmetic; use code execution for precise math

Q: What does LLM stand for and what does it do?

LLM stands for Large Language Model. It is a neural network trained on massive amounts of text data that generates text by predicting the next word in a sequence. By repeating this prediction step-by-step, it can write essays, answer questions, summarize documents, and hold conversations.

Q: What is the difference between a traditional ML model and an LLM?

Traditional ML models are trained for specific tasks (like spam classification or sentiment analysis) and require labeled data and feature engineering for each task. LLMs are general-purpose: one single model trained on diverse internet text can perform hundreds of different tasks — translation, summarization, code generation, Q&A, creative writing — without needing separate training for each one.

Q: Why are LLMs called “large”? What makes them large?

“Large” refers to three things: (1) the number of parameters — modern LLMs have tens to hundreds of billions of parameters, (2) the training data — they are trained on trillions of tokens from the internet, books, and code repositories, and (3) the computational resources — training requires thousands of GPUs running for weeks or months. The combination of these three factors creates emergent abilities that smaller models lack.

Q: Explain why LLMs appear intelligent but are not actually thinking.

LLMs appear intelligent because they have been trained on virtually every text ever written by humans. When asked a question, the model doesn’t “think” about the answer — it calculates the most likely sequence of tokens based on statistical patterns in its training data. A useful analogy: a LLM is like a parrot that has heard every conversation ever recorded. It can say the right thing in the right context because it has heard it so many times, not because it understands what it is saying. This is why LLMs can confidently state falsehoods (hallucinations) — they are producing statistically plausible text, not verifying truth.

Q: What are emergent abilities in LLMs, and why do they matter?

Emergent abilities are capabilities that appear only when a model reaches a certain size threshold — they are not present in smaller versions of the same architecture. Examples include: chain-of-thought reasoning, in-context learning (learning a new task from examples without weight updates), code generation, and solving arithmetic problems. These abilities are notoriously unpredictable — researchers don’t know exactly at what scale they will appear or which abilities will emerge. This is significant because it means our understanding of LLM capabilities is always incomplete: a model 2x larger may do things we didn’t expect, and a model 2x smaller may lack abilities we assumed were basic.

Q: What is the difference between a foundation model and a fine-tuned model?

A foundation model (e.g., GPT-3 base, LLaMA-3 base) is trained on raw internet text using self-supervised learning — next-token prediction. It is extremely capable but hard to use because it continues text rather than following instructions. A fine-tuned model (e.g., GPT-3.5-turbo, LLaMA-3-Chat) is a foundation model that has been further trained on instruction-response pairs and preference data (RLHF). Fine-tuning aligns the model to be helpful, follow instructions, and refuse harmful requests. The foundation model is a raw intelligence engine; the fine-tuned model is a polished assistant. Most users interact with fine-tuned models.


ConceptKey Point
LLMLarge Language Model — neural network trained on internet-scale text to predict the next token
”Large”Billions of parameters, trillions of training tokens, thousands of GPUs
”Language”Processes and generates human language text
”Model”A mathematical approximation of language patterns — a prediction engine, not a thinking machine
Key technologyTransformer architecture with self-attention
How it worksTokenize → predict next token → append → repeat
Emergent abilitiesCapabilities that appear at scale (reasoning, in-context learning)
Major familiesGPT (OpenAI), Claude (Anthropic), Gemini (Google), LLaMA (Meta), Mistral, Qwen, DeepSeek
Not thinkingLLMs do NOT understand, believe, or know — they generate statistically plausible text
HallucinationConfident falsehoods — the model generates plausible but incorrect information

Previous: Phase 3 — Deep Learning Cheat Sheet

Next: 02 — How Language Models Work

Related Topics:

Practice Questions:

  1. List 5 different LLMs and describe how they differ
  2. Explain why an LLM is a “prediction engine” rather than a “thinking machine”
  3. What emergent ability of LLMs is most surprising to you and why?
  4. Draw the high-level flow of how an LLM generates a response
  5. Compare a traditional ML spam filter with an LLM — what can the LLM do that the spam filter cannot?

Further Reading: