Performance: profiling, vectorising and when to leave Python
PY · Chapter 611 min readAsked at Two Sigma, Citadel Securities, Jump, Hudson River Trading
After this lesson you should be able to
- Profile before optimising, and know which tool answers which question.
- Rank the standard speedups by how much they buy.
- Say when Python is the wrong language and what replaces it.
Python is slow per operation and fast to write, which makes it right for research and wrong for a hot path. Making research code fast is almost entirely about moving the loop out of the interpreter; the decision to leave Python altogether is a different question with a much higher bar.
Proposition 6.1
Measure before you optimise
Intuition about where time goes is unreliable, and the cost of being wrong is optimising code that was never the bottleneck. Profile, find the dominant cost, fix that, and profile again — because the second bottleneck is rarely where you expected after the first is gone.
Holds when
cProfilefor function-level costs;line_profilerwhen you need the line.memory_profilerortracemallocwhen the problem is memory rather than time.%timeitfor microbenchmarks, and beware of caching effects that make the second run unrepresentative.
| Step | Typical gain | Cost to you |
|---|---|---|
| Fix the algorithm | Unbounded | Thinking |
| Vectorise with NumPy | 10–100 | Rewriting the loop |
| Avoid repeated allocation | 2–5 | Preallocate and use out= |
numba on a numeric loop | 10–100 | A decorator, and a compile step |
| Multiprocessing | Number of cores | Serialisation overhead |
| Rewrite in C++ | 2–10 over good NumPy | A great deal |
The rest of this lesson is in Premium
You have read the opening. 10 more sections follow, including 3 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read Complexity: reading it off, and deriving it in full, free.