ExpertMultiple choice
A rolling mean, faster · Part 3 of 3
You compute a 50-day rolling mean over a list p of 1,000 prices, producing one value for each full window. Count one read each time a price is fetched from the list.
To vectorise it you take c = np.cumsum(p) and form each window sum as c[i] - c[i - 50], on ten million prices near 10,000. The late windows come out slightly wrong. Why?
- ALate cumulative sums are near , where their rounding error is far larger than a window sum can absorb
- B
np.cumsumuses pairwise summation, which is less accurate than a plain loop - CThe cumulative sum overflows once it passes about two billion, because NumPy accumulates into a 32-bit integer by default
- DThe windows are misaligned by one element
The worked solution is in Premium
The answer, the full working and the one idea to take away – for this and all 1,322 questions in the bank. Answer it in practice and your working is marked, with a known mistake named when you make one.
Learn the method
Performance: profiling, vectorising and when to leave Python
More python and data for quants questions
- You test x in container millions of times against a fixed collection of 100,000 ids.Foundation
- Profiling shows 60% of runtime in one function.Applied
- A NumPy array is C-ordered. Which loop is faster: summing along rows or along columns?Advanced
- You add an array of shape (100, 5) to one of shape (5,). What happens?Foundation
- What happens with def f(x, acc=[]): acc.append(x); return acc?Foundation
- How much memory does a 5,000 × 5,000 float64 NumPy array use, in megabytes?Foundation