Skip to content

The Global Interpreter Lock (GIL)

The Global Interpreter Lock (GIL) is a mutex that prevents multiple threads from executing Python bytecode simultaneously. It ensures thread safety but limits parallelism for CPU-bound tasks.

import sys
# The GIL ensures that only one thread executes
# Python bytecode at a time — even on multi-core CPUs
# CPU-bound task — GIL prevents parallelism
def count_numbers():
total = 0
for i in range(10_000_000):
total += i
return total
flowchart TB
subgraph Threading["Threading (Same Process — GIL)"]
direction TB
T1["Thread 1"]
T2["Thread 2"]
T3["Thread 3"]
L1["GIL Lock 🚫"]
T1 -->|waits| L1
T2 -->|waits| L1
T3 -->|waits| L1
L1 -->|"one at a time"| CPU1["1 CPU Core"]
end
subgraph Multiprocessing["Multiprocessing (Separate Processes)"]
direction TB
P1["Process 1"]
P2["Process 2"]
P3["Process 3"]
G1["GIL 1"]
G2["GIL 2"]
G3["GIL 3"]
P1 --> G1 --> C1["CPU Core 1"]
P2 --> G2 --> C2["CPU Core 2"]
P3 --> G3 --> C3["CPU Core 3"]
end
style Threading fill:#1e40af,color:#fff
style Multiprocessing fill:#7c3aed,color:#fff
style L1 fill:#dc2626,color:#fff
style CPU1 fill:#6b7280,color:#fff
style G1 fill:#059669,color:#fff
style G2 fill:#059669,color:#fff
style G3 fill:#059669,color:#fff
style C1 fill:#2563eb,color:#fff
style C2 fill:#2563eb,color:#fff
style C3 fill:#2563eb,color:#fff
import time
import threading
import multiprocessing
def cpu_intensive(n):
"""Heavy CPU computation"""
return sum(i * i for i in range(n))
# Threading — NO speedup for CPU tasks (GIL)
start = time.time()
threads = [threading.Thread(target=cpu_intensive, args=(10_000_000,))
for _ in range(4)]
for t in threads: t.start()
for t in threads: t.join()
print(f"Threading: {time.time() - start:.2f}s") # No faster than single thread!
# Multiprocessing — YES speedup (separate processes, separate GILs)
start = time.time()
with multiprocessing.Pool(4) as pool:
pool.map(cpu_intensive, [10_000_000] * 4)
print(f"Multiprocessing: {time.time() - start:.2f}s") # ~4x faster!
# I/O-bound tasks — GIL is released during I/O operations
import requests
def fetch_url(url):
response = requests.get(url) # GIL released during I/O wait
return response.status_code
# Threading works great here!
urls = ["http://example.com"] * 20
with ThreadPoolExecutor(max_workers=10) as executor:
results = list(executor.map(fetch_url, urls))
# C extensions release the GIL
import numpy as np
# numpy operations run without the GIL — true parallelism!
  1. Use multiprocessing for CPU-bound tasks
  2. Use C extensions (numpy, pandas, numba) that release the GIL
  3. Use asyncio for high-concurrency I/O
  4. Use JIT compilers like PyPy (no GIL in some implementations)
  5. Use Python 3.12+ — improved GIL behavior with sub-interpreters
  6. Use nogil (experimental Python fork without GIL)
import sys
print(f"Python {sys.version}")
# Python 3.12 introduced per-interpreter GIL
# Support for sub-interpreters with separate GILs
import _xxsubinterpreters as interpreters
interp = interpreters.create()
interpreters.run_string(interp, "import this")
  1. Don’t blame the GIL until you’ve measured your bottleneck
  2. Use threading for I/O, multiprocessing for CPU
  3. Use C extensions for heavy number crunching
  4. Profile before optimizing — the GIL may not be your actual problem
  5. Consider alternative Python implementations if GIL is a true blocker