CPU Profiling: cProfile, timeit & py-spy

Before attempting code optimizations, you must measure runtime execution using empirical profiling tools. Understanding the difference between Deterministic Profiling (cProfile), Micro-benchmarking (timeit), and Sampling Profiling (py-spy) is critical for identifying real CPU bottlenecks without guessing or introducing measurement overhead.

This chapter details deterministic vs sampling profiling, cProfile analysis via pstats, timeit garbage collection isolation, and zero-overhead production sampling with py-spy.


1. Deterministic vs. Sampling Profilers

CPU Profiling Methodologies:

1. DETERMINISTIC PROFILING (cProfile):
   Injects C hooks on EVERY function entry/exit.
   - Pros: Exact call counts, 100% deterministic measurements.
   - Cons: High measurement overhead (2x-5x slowdown); alters runtime behavior.

2. SAMPLING PROFILING (py-spy):
   Inspects CPython process memory out-of-band at 100Hz intervals.
   - Pros: ZERO overhead (~1%-2%); safe for live production servers!
   - Cons: Statistical sampling (does not capture sub-millisecond call counts).

2. Deterministic CPU Profiling with cProfile & pstats

cProfile is Python’s standard deterministic profiler written in C. It tracks total execution time (tottime), cumulative time (cumtime), and call counts (ncalls):

# Run script under cProfile and output binary statistics file
python -m cProfile -o profile.prof my_script.py

Analyzing Statistics with pstats / snakeviz:

import pstats

p = pstats.Stats("profile.prof")
p.strip_dirs().sort_stats("cumulative").print_stats(10) # Print top 10 functions by cumulative time!
  • tottime: Total time spent inside the target function excluding time spent calling sub-functions.
  • cumtime: Cumulative time spent inside the function including all sub-function calls.

3. Production Sampling Profiling with py-spy

Running cProfile on a production web server degrades request throughput by 50%+.

py-spy is a sampling profiler written in Rust that reads CPython process memory directly using OS process_vm_readv(2) system calls without modifying target process bytecode or injecting hooks:

# Profile running production Gunicorn process (PID 1234) and generate Flamegraph SVG
py-spy record --pid 1234 --output flamegraph.svg --rate 100
Flamegraph Visualization Architecture:

[ Wide Base Rectangles: High-level caller functions (main/handler) ]
                             |
                             v (Stack depth grows vertically)
[ Narrow Top Rectangles: Leaf functions consuming raw CPU cycles ]

4. Micro-benchmarking with timeit

Never benchmark code snippets using time.time(). OS scheduling pauses and garbage collection sweeps pollute raw time delta measurements.

The timeit module runs code snippets millions of times, temporarily disabling Garbage Collection to measure pure CPU execution time:

import timeit

# Micro-benchmark list comprehension vs map()
t1 = timeit.timeit("[x * 2 for x in range(1000)]", number=10000)
t2 = timeit.timeit("list(map(lambda x: x * 2, range(1000)))", number=10000)

print(f"Comprehension: {t1:.4f}s | Map: {t2:.4f}s")
Display Options
Appearance
Text Size
100%