09. Loss Functions
Introduction
Section titled “Introduction”A loss function measures how wrong the network’s prediction is. The goal of training is to make this number as small as possible.
Think of loss as a report card — it tells the network exactly how far off it is, and in which direction to improve.
The Role of Loss in Learning
Section titled “The Role of Loss in Learning”flowchart LR Input["Input Data"] --> Model["Neural Network"] --> Pred["Prediction ŷ"] Truth["True Label y"] --> Loss["Loss Function\nL(y, ŷ)"] Pred --> Loss Loss --> BackProp["Backpropagation\n(use loss to update weights)"] BackProp --> Model
style Loss fill:#ef4444,color:#fffWithout a loss function, the network has no signal for learning. The loss is the compass.
Real-World Analogy
Section titled “Real-World Analogy”You’re learning to throw darts at a bullseye:
- Prediction (ŷ) = where your dart landed
- True label (y) = the bullseye
- Loss = how far your dart is from the bullseye
- Training = gradually adjusting your throw to minimize that distance
Every throw, you measure the error. You adjust your technique based on that measurement. Over many throws, you improve.
graph LR Throw["Throw dart"] --> Measure["Measure distance from bullseye\n= Loss"] --> Adjust["Adjust technique\n= Update weights"] --> ThrowLoss Function 1: Mean Squared Error (MSE)
Section titled “Loss Function 1: Mean Squared Error (MSE)”Used for regression tasks — predicting continuous values.
MSE = (1/n) × Σ(y - ŷ)²Intuition: Average squared distance between predictions and true values.
Example:
| Predicted house price | Actual price | Error | Squared Error |
|---|---|---|---|
| $450,000 | $400,000 | $50,000 | 2,500,000,000 |
| $320,000 | $350,000 | -$30,000 | 900,000,000 |
| $280,000 | $300,000 | -$20,000 | 400,000,000 |
MSE = (2.5B + 0.9B + 0.4B) / 3 = $1.27B
Why squared?
- Always positive (can’t cancel out)
- Penalizes large errors more heavily than small ones
graph LR subgraph MSE_Plot["Loss vs Error"] A["Small error = small loss"] B["Large error = VERY large loss\n(squared effect)"] endimport numpy as npimport tensorflow as tf
# Manual MSEy_true = np.array([400000, 350000, 300000])y_pred = np.array([450000, 320000, 280000])
mse = np.mean((y_true - y_pred) ** 2)print(f"MSE: {mse:,.0f}")
# Kerasmodel.compile(optimizer='adam', loss='mse')
# For regression tasks:model = tf.keras.Sequential([ tf.keras.layers.Dense(64, activation='relu'), tf.keras.layers.Dense(1) # No activation for regression])model.compile(loss='mean_squared_error', optimizer='adam')Loss Function 2: Mean Absolute Error (MAE)
Section titled “Loss Function 2: Mean Absolute Error (MAE)”MAE = (1/n) × Σ|y - ŷ|Intuition: Average absolute distance between predictions and true values.
MSE vs MAE:
| MSE | MAE | |
|---|---|---|
| Sensitivity to outliers | High (squares large errors) | Low (linear penalty) |
| Optimization | Smooth, easy to optimize | Less smooth (not differentiable at 0) |
| Interpretation | Harder to interpret (squared units) | Easy: in original units |
| Use when | Outliers are important to penalize | Outliers are noise to ignore |
# Manual MAEmae = np.mean(np.abs(y_true - y_pred))print(f"MAE: {mae:,.0f}")
# Kerasmodel.compile(loss='mean_absolute_error', optimizer='adam')Loss Function 3: Binary Cross-Entropy
Section titled “Loss Function 3: Binary Cross-Entropy”Used for binary classification (spam/not spam, cat/dog, yes/no).
BCE = -(1/n) × Σ[y·log(ŷ) + (1-y)·log(1-ŷ)]Intuition: Measures how confident the prediction was and whether that confidence was correct.
graph LR subgraph Cases["What the formula captures"] A["y=1, ŷ=0.99\n→ Low loss ✓\n(Confident and correct)"] B["y=1, ŷ=0.50\n→ Medium loss\n(Uncertain)"] C["y=1, ŷ=0.01\n→ Very high loss ✗\n(Confident but WRONG)"] endWhy logarithm? log(x) approaches -∞ as x approaches 0, so being confidently wrong is severely penalized.
# Manual Binary Cross-Entropydef binary_cross_entropy(y_true, y_pred, epsilon=1e-15): y_pred = np.clip(y_pred, epsilon, 1 - epsilon) # Avoid log(0) return -np.mean(y_true * np.log(y_pred) + (1 - y_true) * np.log(1 - y_pred))
# Example: email spam detectiony_true = np.array([1, 0, 1, 1, 0]) # Spam (1) or not (0)y_pred = np.array([0.9, 0.1, 0.8, 0.3, 0.2]) # Model confidence
loss = binary_cross_entropy(y_true, y_pred)print(f"Binary Cross-Entropy: {loss:.4f}")
# Kerasmodel.compile( loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])# Output layer: Dense(1, activation='sigmoid')Loss Function 4: Categorical Cross-Entropy
Section titled “Loss Function 4: Categorical Cross-Entropy”Used for multi-class classification (10 digits, 1000 ImageNet classes).
CCE = -(1/n) × Σ Σ yᵢⱼ · log(ŷᵢⱼ)Intuition: Like binary cross-entropy but for multiple classes. Penalizes wrong predictions based on confidence.
# Example: MNIST digit classificationy_true = np.array([[0, 0, 1, 0, 0, 0, 0, 0, 0, 0]]) # True class: 2y_pred = np.array([[0.01, 0.02, 0.85, 0.04, 0.02, 0.01, 0.01, 0.01, 0.02, 0.01]])
cce = -np.sum(y_true * np.log(y_pred + 1e-15))print(f"CCE: {cce:.4f}") # Low loss — model was confident and correct
# Kerasmodel.compile( loss='categorical_crossentropy', # Labels are one-hot encoded # or: loss='sparse_categorical_crossentropy', # Labels are integers (0-9) optimizer='adam', metrics=['accuracy'])# Output layer: Dense(10, activation='softmax')Choosing the Right Loss Function
Section titled “Choosing the Right Loss Function”flowchart TD A["What is your task?"] --> B{"Regression\n(predict a number)?"} B -- Yes --> C{"Outliers\nimportant?"} C -- Yes --> D["Mean Squared Error\n(MSE)"] C -- No --> E["Mean Absolute Error\n(MAE)"]
B -- No --> F{"Classification?"} F -- Binary --> G["Binary Cross-Entropy\n(2 classes, sigmoid output)"] F -- Multi-class --> H{"Labels format?"} H -- "One-hot [0,0,1,0]" --> I["Categorical Cross-Entropy"] H -- "Integer (2)" --> J["Sparse Categorical CE"]Tracking Loss During Training
Section titled “Tracking Loss During Training”import tensorflow as tfimport numpy as np
# Build modelmodel = tf.keras.Sequential([ tf.keras.layers.Dense(64, activation='relu', input_shape=(10,)), tf.keras.layers.Dense(32, activation='relu'), tf.keras.layers.Dense(1, activation='sigmoid')])model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
# Generate dummy dataX = np.random.randn(1000, 10)y = (X[:, 0] + X[:, 1] > 0).astype(float)
# Train — watch loss decreasehistory = model.fit(X, y, epochs=10, validation_split=0.2)
# Plot lossfor epoch, (loss, val_loss) in enumerate( zip(history.history['loss'], history.history['val_loss'])): print(f"Epoch {epoch+1}: loss={loss:.4f}, val_loss={val_loss:.4f}")JavaScript: Custom Loss Visualization
Section titled “JavaScript: Custom Loss Visualization”import * as tf from '@tensorflow/tfjs';
// Binary cross-entropy from scratchfunction binaryCrossEntropy(yTrue, yPred) { const epsilon = 1e-7; const clipped = yPred.map(p => Math.min(Math.max(p, epsilon), 1 - epsilon)); const losses = yTrue.map((y, i) => -(y * Math.log(clipped[i]) + (1 - y) * Math.log(1 - clipped[i])) ); return losses.reduce((sum, l) => sum + l, 0) / losses.length;}
// Confidence vs loss visualizationconst scenarios = [ { label: 'Correct + Confident', y: 1, pred: 0.99 }, { label: 'Correct + Uncertain', y: 1, pred: 0.50 }, { label: 'Wrong + Uncertain', y: 1, pred: 0.20 }, { label: 'Wrong + Confident', y: 1, pred: 0.01 },];
scenarios.forEach(({ label, y, pred }) => { const loss = binaryCrossEntropy([y], [pred]); console.log(`${label}: loss = ${loss.toFixed(4)}`);});// Correct + Confident: loss = 0.0101 (low)// Correct + Uncertain: loss = 0.6931// Wrong + Uncertain: loss = 1.6094// Wrong + Confident: loss = 4.6052 (very high!)Interview Questions
Section titled “Interview Questions”Q1: What is a loss function?
A loss function quantifies the difference between the model’s predictions and the true labels. It produces a scalar value that the optimizer tries to minimize through training. Common examples: MSE for regression, cross-entropy for classification.
Q2: Why use cross-entropy instead of MSE for classification?
MSE treats all errors linearly. Cross-entropy uses log-probability, which heavily penalizes confident wrong predictions. Being 99% confident and wrong is astronomically worse than being 51% confident and wrong — cross-entropy captures this properly. MSE doesn’t. Additionally, cross-entropy produces gradients that are more informative for classification tasks.
Q3: What happens if loss doesn’t decrease during training?
Possible causes: (1) Learning rate too high — gradients diverge. (2) Learning rate too low — extremely slow convergence. (3) Wrong loss function for the task. (4) Model capacity too low (underfitting). (5) Data/label errors. (6) Weight initialization issues.
Q4: What’s the difference between categorical_crossentropy and sparse_categorical_crossentropy?
Both are the same mathematical loss, but differ in expected label format.
categorical_crossentropyexpects one-hot encoded labels:[0, 0, 1, 0, 0].sparse_categorical_crossentropyexpects integer class indices:2. Use sparse when you have many classes to save memory.
Best Practices
Section titled “Best Practices”- Match loss to task — Regression → MSE/MAE, binary classification → BCE, multi-class → CCE
- Monitor validation loss — Training loss always decreases; validation loss tells you about generalization
- Use
sparse_categorical_crossentropywhen labels are integers (most common in practice) - Watch for NaN loss — Usually caused by exploding gradients or log(0); add gradient clipping or check preprocessing
- Loss ≠ Accuracy — Loss is differentiable and used for optimization; accuracy is human-interpretable but not directly optimized
Common Mistakes
Section titled “Common Mistakes”- Using MSE for classification — Produces poor gradients; always use cross-entropy for classification
- Not normalizing inputs — Can cause loss to start extremely high and converge slowly
- Ignoring validation loss — Overfitting shows as training loss decreasing while validation loss increases
- Forgetting
from_logits=True— When using softmax inside the model, Keras handles it; when you don’t have softmax in the model but want cross-entropy, setfrom_logits=True
Summary
Section titled “Summary”| Loss Function | Task | Output Layer | When to Use |
|---|---|---|---|
| MSE | Regression | Linear | Continuous target, outliers matter |
| MAE | Regression | Linear | Continuous target, robust to outliers |
| Binary CE | Binary classification | Sigmoid (1 unit) | Yes/No, Spam/Ham |
| Categorical CE | Multi-class | Softmax (N units) | One-hot labels |
| Sparse Categorical CE | Multi-class | Softmax (N units) | Integer labels |
Navigation
Section titled “Navigation”Previous: 08 — Activation Functions
Next: 10 — Backpropagation
Related Topics:
Practice Exercises
Section titled “Practice Exercises”- Compute BCE loss manually for a batch of 5 predictions
- Train two identical networks — one with MSE, one with cross-entropy for classification — compare results
- Deliberately create a “wrong + confident” prediction — observe the large loss value
- Plot training vs validation loss curves — identify when overfitting begins
- What does a loss of 0.0 mean? Is that good?