Skip to content

08. Common AI Terminology

AI has a dense vocabulary. The same concept sometimes has multiple names. This glossary gives you the precise definition for each term and how they relate.


TermDefinition
ModelA mathematical function mapping inputs to outputs, learned from data
Parameters / WeightsNumerical values inside a model adjusted during training
TrainingThe process of adjusting weights to minimize prediction error
InferenceUsing a trained model to make predictions on new data
DatasetCollection of examples used to train, validate, or test a model
FeatureAn input variable used by the model (e.g., age, price, word)
Label / TargetThe correct output the model should predict
Loss / CostA number measuring how wrong the model’s prediction is
Gradient DescentOptimization algorithm that updates weights to reduce loss
BackpropagationAlgorithm to compute how each weight contributes to the loss
EpochOne complete pass through the training dataset
Batch SizeNumber of examples processed in one training step
Learning RateHow large each weight update step is

TermDefinition
OverfittingModel memorizes training data, poor on new data
UnderfittingModel too simple, poor on both training and new data
GeneralizationModel performs well on unseen data
Bias (statistical)Systematic error — model consistently wrong in one direction
VarianceModel too sensitive to training data — inconsistent across samples
RegularizationTechniques to prevent overfitting (L1, L2, dropout)
Cross-validationEvaluating model on multiple train/test splits to reduce variance
BaselineSimple model to beat — sets minimum acceptable performance

TermDefinition
Neural NetworkLayers of interconnected nodes loosely inspired by the brain
Neuron / NodeSingle unit that computes a weighted sum + activation
LayerGroup of neurons (input, hidden, output layers)
Activation FunctionNon-linear function applied per neuron (ReLU, sigmoid, tanh)
Deep LearningNeural networks with many hidden layers
CNNConvolutional Neural Network — specialized for images
RNNRecurrent Neural Network — specialized for sequences
TransformerArchitecture using self-attention for parallel sequence processing
AttentionMechanism that weighs the importance of different input parts
EmbeddingDense vector representation of discrete items (words, images)

TermDefinition
TokenUnit of text the model processes (word, sub-word, or character)
Context WindowMaximum tokens a model can process at once
PromptInput text given to an LLM
Completion / ResponseModel’s output
TemperatureControls randomness of output (0 = deterministic, 1+ = creative)
Top-p / Top-kSampling strategies controlling output diversity
HallucinationModel confidently generates false information
Fine-tuningFurther training a pretrained model on domain-specific data
RLHFReinforcement Learning from Human Feedback — used to align LLMs
RAGRetrieval-Augmented Generation — grounding LLM with external docs

TermDefinition
Supervised LearningTraining with labeled examples (input → correct output)
Unsupervised LearningFinding patterns without labels
Reinforcement LearningLearning via reward/penalty signals
Train SetData used to fit the model
Validation SetData used to tune hyperparameters
Test SetData held out for final evaluation only
Data DriftDistribution of input data changes over time
Class ImbalanceOne output class much rarer than others (e.g., fraud)

Q: What is the difference between a parameter and a hyperparameter?

A: Parameters (weights) are learned from training data — they are what the model optimizes. Hyperparameters are settings you choose before training — like learning rate, batch size, number of layers, or dropout rate. Parameters are internal to the model; hyperparameters are external knobs you tune.


Q: What is the difference between overfitting and underfitting?

A: Overfitting happens when a model learns training data too well — including noise — and fails to generalize. It has low training error but high test error. Underfitting happens when a model is too simple to capture the real patterns — high error on both training and test data. The goal is the sweet spot between them (good generalization).