Why Python is slow — and how to mitigate it
CPython prioritizes simplicity, dynamism, and readability over raw speed. The cost is real: pure-Python compute is typically 10-100x slower than equivalent C/Rust/Go.
Where the slowness comes from
1. Interpretation overhead
CPython evaluates bytecode in a Python-level loop. Every operation goes through a dispatch step.
A single a + b requires: load a, load b, look up the + method on a’s type, call it, possibly try the reverse method on b, push result. In C, int + int is one instruction.
2. Boxing
Every Python value is a heap-allocated object with a refcount, type pointer, and value. 1 isn’t 4 bytes — it’s ~28 bytes. Iterating an int list dereferences a pointer per element.
In NumPy, np.array([1, 2, 3], dtype=np.int32) is 12 bytes — a contiguous C array. That’s why NumPy is 100x faster on numeric loops.
3. Dynamic dispatch
Every method call is a dict lookup on the type’s MRO. JIT compilers (PyPy) cache these to skip the lookup; CPython pays the lookup every call (3.11+ has specialized adaptive opcodes that mitigate this somewhat).
4. The GIL
Only one Python thread runs Python bytecode at a time, regardless of CPU count. CPU-bound multi-threaded Python doesn’t scale beyond 1 core. See 04_async_concurrency/01_gil.md.
5. Garbage collection / refcounting
Every assignment / deassignment touches refcounts. Tens of millions of refcount ops per second add overhead to allocation-heavy code.
How to make it fast
In rough order from cheapest to most invasive:
1. Use the right algorithm and data structure
A pure-Python O(n log n) solution beats a C-extension O(n²) for any n above a few thousand. Profile before optimizing constants.
2. Use built-ins and stdlib
sum(iterable) is a C loop. [x*2 for x in big] is faster than for x in big: result.append(x*2) because the comprehension has a single bytecode opcode for append.
map, filter, sorted, min, max, any, all are all C-implemented loops.
3. Avoid recomputation
@lru_cache, @cached_property. Pull invariants out of loops (loop-invariant code motion):
# slow
for x in big_list:
result.append(some_obj.method(x))
# faster
m = some_obj.method # bound method lookup once
for x in big_list:
result.append(m(x))
4. Use NumPy / pandas for numeric work
Replace Python loops with vectorized array operations:
# slow
result = [x ** 2 + 1 for x in arr] # ~1 sec for 10M items
# fast
import numpy as np
arr = np.array(arr)
result = arr ** 2 + 1 # ~30 ms
Vectorized ops run in C with no boxing.
5. Release the GIL with C extensions / native libs
NumPy, Pandas, scikit-learn, PyTorch, etc., release the GIL during their C-level computation. Threading does help when most work is in those libraries.
6. Use multiprocessing for CPU parallelism
For pure-Python CPU work, fork separate processes:
from concurrent.futures import ProcessPoolExecutor
with ProcessPoolExecutor() as pool:
results = list(pool.map(heavy_compute, chunks))
Cost: pickling overhead crossing process boundaries, no shared memory by default.
7. Use Cython, Numba, mypyc, or Rust
For pure-Python hot loops:
- Cython — compile annotated Python to C. Mature; tooling is heavy.
- Numba —
@jitdecorator; compiles numeric Python to native via LLVM. Great for numerical hot loops. - mypyc — compile type-annotated Python to C. Used by mypy itself.
- Rust extensions (PyO3) — write a Rust module, expose to Python. Best for new performance-critical libraries.
8. Try PyPy
PyPy is a JIT-compiled Python implementation. 5-50x faster on pure-Python workloads. Drawback: smaller ecosystem, slower NumPy/pandas, slower startup.
9. Profile-guided optimization is faster than guessing
Always profile (see 01_profiling_basics.md). Hot paths are rarely where you expect.
Numbers to know (rough orders of magnitude)
- Pure Python loop: ~10-100M ops/sec
- NumPy vectorized: ~1B ops/sec (single-core)
- Dict lookup: ~100ns
- Function call overhead: ~50-100ns
- Attribute lookup: ~30ns
- Local variable read: ~5ns
- Network roundtrip (LAN): ~0.5ms = 5M Python ops worth
That last point: the difference between “slow Python” and “fast Python” is usually drowned out by I/O. Optimize Python only when you’ve already minimized I/O and database calls.
Interview angle
- “Why is Python slower than C?” (Interpretation, boxing, dynamic dispatch, refcount overhead, GIL.)
- “How would you speed up a CPU-bound function?” (Profile first → vectorize with NumPy → multiprocessing → Numba/Cython → Rust extension.)
- “When does threading help in Python?” (I/O-bound work, or when most CPU is in GIL-releasing native libs like NumPy.)