Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • MKTMarkets and products
    • CSData structures and algorithms
    • PYPython and data for quants
      • 1Python fundamentals

        • Idiomatic Python: generators, itertools and the gotchas
      • 2NumPy

        • NumPy: broadcasting, axes and vectorisation
      • 3pandas

        • pandas: groupby, joins, resampling and time zones
      • 4Market data handling

        • Market data in pandas, and where look-ahead hides
      • 5Vectorised backtesting

        • Vectorised backtesting: signal to position to P&L
      • 6Performance

        • Performance: profiling, vectorising and when to leave Python
    • NUMNumerical methods
    • SYSSystems and low latency

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative development
  3. /Python and data for quants
  4. /Vectorised backtesting

Vectorised backtesting: signal to position to P&L

PY · Chapter 5·12 min read·Asked at Two Sigma, Citadel Securities, Point72, AQR

Assumes pandas: groupby, joins, resampling and time zones.

After this lesson you should be able to

  • Build the signal-to-P&L pipeline with correct alignment.
  • Account for costs and turnover properly.
  • Compute the standard performance statistics without error.

A backtest is four arrays — signal, position, return, cost — and the entire difficulty is in the alignment between them. Written vectorised, the lag is a visible line of code that a reviewer can check; written as a loop, it is an assumption nobody can see.

StageOperationThe trap
SignalCross-sectional z-score per dateFull-sample normalisation
PositionScale to a risk target, then lagForgetting the lag
Gross P&LPosition times forward returnUsing the same-period return
TurnoverAbsolute change in positionIgnoring it entirely
CostsTurnover times a cost modelA fixed cost per share
Net P&LGross minus costsReporting only the gross
Table 5.1 · The pipeline. Every row has a trap and every trap inflates the result. A backtest that has not been checked against all six is optimistic by default, not by accident.
import numpy as np
import pandas as pd

def backtest(signal: pd.DataFrame, returns: pd.DataFrame,
             cost_bps: float = 5.0, target_vol: float = 0.10) -> pd.Series:
    """signal and returns are (dates x assets), aligned and already lagged-safe."""
    # 1. Cross-sectional z-score, within each date only.
    z = signal.sub(signal.mean(axis=1), axis=0).div(signal.std(axis=1), axis=0)

    # 2. Dollar-neutral weights summing to one unit of gross exposure.
    w = z.div(z.abs().sum(axis=1), axis=0)

    # 3. Act tomorrow on what you knew today.
    pos = w.shift(1)

    gross = (pos * returns).sum(axis=1)
    turnover = pos.diff().abs().sum(axis=1)
    cost = turnover * cost_bps / 1e4

    net = gross - cost
    return net * (target_vol / (net.std() * np.sqrt(252)))
Listing 5.2 · The whole backtest. The shift(1) on line 3 is the entire difference between a backtest and a fantasy. Keeping it on its own line, with a comment, is deliberate — it is the line a reviewer should be able to find in two seconds. Time O(dates x assets) · Space O(dates x assets).

How much to lag. One bar is the minimum and rarely the truth. A signal computed from the close cannot be traded at that close, so the honest lag is to the next open at least — and if the data arrives with a publication delay, the lag is that delay rather than one bar. The right question is not "how many periods should I shift" but "when could I first have acted on this", and the answer comes from the data’s release schedule, not from the frequency of the index.

StatisticFormulaThe error people make
Annualised returnMean daily ×252\times 252×252Mixing arithmetic and geometric
Annualised volatilityDaily std ×252\times \sqrt{252}×252​Multiplying by 252
Sharpe ratioAnnualised excess return over volatilityForgetting the risk-free rate
Max drawdownLargest peak-to-trough on the cumulative curveComputing it on returns rather than levels
TurnoverSum of absolute position changesCounting one side only
Table 5.3 · The statistics, computed correctly. The volatility row is the one that appears in real code. Annualising a standard deviation by multiplying rather than taking the square root inflates it by a factor of nearly sixteen, and the resulting Sharpe is comically small rather than obviously wrong.

The rest of this lesson is in Premium

You have read the opening. 10 more sections follow, including 5 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read Complexity: reading it off, and deriving it in full, free.

← Market data in pandas, and where look-ahead hidesPerformance: profiling, vectorising and when to leave Python →
On this page
  • The pipeline
  • The whole backtest
  • The statistics, computed correctly

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.