Skip to content

15. AI Cheat Sheet

┌─────────────────────────────────────────────┐
│ AI │
│ (any machine simulating human intelligence) │
│ ┌───────────────────────────────────────┐ │
│ │ Machine Learning │ │
│ │ (AI that learns from data) │ │
│ │ ┌─────────────────────────────────┐ │ │
│ │ │ Deep Learning │ │ │
│ │ │ (ML with neural networks) │ │ │
│ │ │ ┌───────────────────────────┐ │ │ │
│ │ │ │ Transformers │ │ │ │
│ │ │ │ (attention-based DL) │ │ │ │
│ │ │ │ ┌─────────────────────┐ │ │ │ │
│ │ │ │ │ LLMs │ │ │ │ │
│ │ │ │ └─────────────────────┘ │ │ │ │
│ │ │ └───────────────────────────┘ │ │ │
│ │ └─────────────────────────────────┘ │ │
│ └───────────────────────────────────────┘ │
└─────────────────────────────────────────────┘

ClassificationTypesExists?
By CapabilityNarrow AI✓ Yes
By CapabilityAGI✗ No
By CapabilitySuper AI✗ No
By FunctionalityReactive Machines✓ Yes (Deep Blue)
By FunctionalityLimited Memory✓ Yes (ChatGPT, Tesla)
By FunctionalityTheory of Mind✗ No
By FunctionalitySelf-Aware✗ No

Data → Train (minimize loss via backprop) → Model (frozen weights) → Inference

1. Problem Definition → Is this an ML problem? What's the metric?
2. Data Collection → Labeled/unlabeled, quality > quantity
3. Data Preparation → Clean, label, engineer features, split (60-80% of time)
4. Model Development → Choose arch, train, tune hyperparameters
5. Evaluation → Test set, fairness, latency, cost
6. Deployment → API, A/B test, rollback plan
7. Monitor & Maintain → Watch for data drift, retrain as needed

TermOne-Line Definition
ModelLearned function: input → output
Parameters/WeightsNumbers adjusted during training
TrainingIterative weight updates to reduce loss
InferenceUsing frozen model to predict
LossHow wrong the prediction is
Gradient DescentOptimizer that walks weights toward lower loss
BackpropagationComputes gradient through the network
EpochOne full pass through training data
OverfittingMemorizes training, fails on new data
UnderfittingToo simple, fails on everything
HallucinationLLM generates confident but false output
TokenUnit of text LLM processes
Context WindowMax tokens LLM can consider at once
RAGGround LLM answers in retrieved documents
Fine-tuningFurther train pretrained model on new data
EmbeddingDense vector representing a piece of content
Data DriftInput distribution changes in production
RLHFAlign model via human preference feedback

Can DoCannot Do (reliably)
Generate fluent textVerify facts
Classify imagesReason about physics
Translate languagesMaintain long-term memory
Detect patterns in dataUnderstand cause vs correlation
Code completionKnow its own knowledge cutoff
Summarize documentsGive consistent answers to edge cases

YearEvent
1950Turing Test proposed
1956”AI” coined at Dartmouth
1997Deep Blue beats Kasparov
2012AlexNet — deep learning for vision
2017Transformer architecture published
2018BERT, GPT-1
2020GPT-3 (175B params)
2022ChatGPT — mainstream AI
2024Reasoning models, AI agents, MCP

PrincipleWhat to Ask
FairnessEqual performance across groups?
TransparencyCan decisions be explained?
PrivacyMinimum data needed? Consent obtained?
AccountabilityWho owns failures?
SafetyDoes it fail safely? Human override?
BeneficenceWho benefits? Who’s harmed?

You now understand:

  • What AI is and where it fits
  • The AI → ML → DL → Transformers → LLMs hierarchy
  • How learning happens (data, loss, backprop, inference)
  • The full AI lifecycle
  • Real-world applications and limitations
  • Ethical considerations

Phase 2: Machine Learning dives into the algorithms themselves — how models actually learn from data, what supervised vs unsupervised vs reinforcement learning looks like in code, and the math behind the patterns.