Skip to content

04. Data & Datasets

In Machine Learning, data is everything. The model can only learn what the data teaches it. Better data beats better algorithms.

Before you can train any model, you need to understand what kind of data you have, what format it’s in, and whether it’s good enough to learn from.


mindmap
root((Data Types))
Structured
CSV
SQL Tables
JSON
Spreadsheets
Unstructured
Images
Audio
Video
Raw Text
Semi-Structured
HTML
XML
Logs
Emails

Rows and columns. Each column is a well-defined attribute.

age,income,education,loan_approved
28,55000,bachelor,1
45,90000,master,1
22,25000,high_school,0

Characteristics:

  • Easy to process with pandas / SQL
  • Most classical ML algorithms work natively on structured data
  • Examples: transaction records, customer databases, sensor readings

No predefined format. Requires transformation before ML can use it.

TypeRaw FormML Representation
ImagePixel values (H×W×3)Flattened array or CNN feature map
TextString of charactersToken IDs, embeddings
AudioSound wave samplesSpectrograms, MFCCs
VideoSequence of image framesFrame embeddings + temporal features

Has some structure but not a strict schema.

{
"user_id": 123,
"events": [
{ "type": "click", "timestamp": "2024-01-01T10:30:00" },
{ "type": "purchase", "amount": 49.99 }
]
}

Requires parsing and transformation to become ML-ready.


Every ML dataset has the same basic structure:

flowchart LR
A[Dataset] --> B[Features / Inputs]
A --> C[Labels / Targets]
B --> D[Column 1: age]
B --> E[Column 2: income]
B --> F[Column 3: city]
C --> G[Column: loan_approved]
TermMeaning
Row / Sample / ExampleOne data point
Column / FeatureOne input attribute
Label / TargetWhat you want to predict
Dataset sizeNumber of rows
DimensionalityNumber of features

A small, clean, representative dataset outperforms a huge, dirty one.

Common data quality problems:

import pandas as pd
df = pd.read_csv("data.csv")
# Check missing values
print(df.isnull().sum())
# Check duplicates
print(f"Duplicates: {df.duplicated().sum()}")
# Check distribution
print(df.describe())
# Check class balance
print(df["label"].value_counts())
ProblemImpactFix
Missing valuesMany algorithms fail or behave poorlyImpute (mean/median/mode) or drop
DuplicatesModel over-learns repeated examplesDrop duplicates
OutliersDistort patternsCap, remove, or transform
Class imbalanceModel predicts majority class alwaysOversample, undersample, or adjust weights
Label noiseModel learns wrong patternsManual review, consensus labeling
BiasModel reflects historical discriminationAudit and diversify data sources

You never train and evaluate on the same data:

flowchart LR
A[Full Dataset 100%] --> B[Train 70-80%]
A --> C[Validation 10-15%]
A --> D[Test 10-15%]
B --> E[Model learns]
C --> F[Hyperparameter tuning]
D --> G[Final evaluation only]
from sklearn.model_selection import train_test_split
# First split: train+val vs test
X_trainval, X_test, y_trainval, y_test = train_test_split(
X, y, test_size=0.15, random_state=42
)
# Second split: train vs validation
X_train, X_val, y_train, y_val = train_test_split(
X_trainval, y_trainval, test_size=0.15, random_state=42
)

There’s no fixed answer, but rough rules of thumb:

TaskMinimum Samples
Binary classification1,000+ per class
Multi-class classification500–1,000+ per class
Regression1,000–10,000+
Image classification (from scratch)10,000+ per class
LLM fine-tuning1,000+ examples

When you don’t have enough data:

  • Transfer learning (start from a pretrained model)
  • Data augmentation (create synthetic variations)
  • Collect more data
  • Use a simpler model

# CSV - most common
df = pd.read_csv("data.csv")
# JSON
import json
with open("data.json") as f:
data = json.load(f)
# SQL Database
import sqlite3
conn = sqlite3.connect("database.db")
df = pd.read_sql("SELECT * FROM customers", conn)
# Parquet (efficient for large datasets)
df = pd.read_parquet("data.parquet")
# Images (PIL)
from PIL import Image
img = Image.open("cat.jpg")
import numpy as np
arr = np.array(img) # shape: (height, width, 3)

DatasetTaskSource
TitanicClassificationKaggle
House PricesRegressionKaggle
MNISTImage classificationTensorFlow/PyTorch
IMDB ReviewsSentiment analysisHugging Face
IrisMulti-class classificationsklearn
Boston HousingRegressionsklearn
# Built-in datasets in sklearn
from sklearn.datasets import load_iris, load_diabetes, load_digits
iris = load_iris()
X, y = iris.data, iris.target
print(X.shape) # (150, 4) — 150 samples, 4 features

Q: Why does data quality matter more than algorithm choice?

A: A model can only learn patterns that exist in training data. If data has missing values, label noise, or doesn’t represent the real-world distribution, no algorithm will fix it. The best algorithm on bad data produces bad predictions. The simplest algorithm on clean, representative data often beats a complex model on noisy data. This is why data preparation takes 60–80% of project time.


Q: What is class imbalance and how do you handle it?

A: Class imbalance occurs when one class is much more frequent than another — for example, 99% legitimate transactions vs 1% fraud. A model trained on this data learns to predict “not fraud” always and achieves 99% accuracy while missing all actual fraud. Solutions: (1) oversample the minority class (SMOTE), (2) undersample the majority class, (3) adjust class weights in the loss function, (4) use precision/recall/F1 instead of accuracy as the metric.


  • Using test data during feature engineering → data leakage
  • Not checking for class imbalance → misleading accuracy
  • Treating the train split as the full dataset → overly optimistic metrics
  • Ignoring temporal ordering in time-series data → future data leaks into training

ConceptKey Point
Structured dataRows and columns — CSV, SQL
Unstructured dataImages, audio, text — needs transformation
Dataset splitsTrain / Validation / Test — never mix
Data qualityMore impactful than algorithm choice
Class imbalanceSpecial handling required for rare events

← Previous: 03. ML Workflow Next →: 05. Features & Labels