Skip to content
QuantMax
QuantMax
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • COMBCounting and combinatorics
    • PROBProbability
    • GAMEGames, decision theory and puzzles
    • MMMarket making
    • MKTMarkets and products

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Formula reference

Probability

9 lessons · 20 equations. Each lesson below gives its formulas and key rules; open the lesson for the full explanation.

Conditional probability and Bayes

Conditional probability

Pr⁡(A∣B)=Pr⁡(A∩B)Pr⁡(B)\Pr(A \mid B) = \frac{\Pr(A \cap B)}{\Pr(B)}Pr(A∣B)=Pr(B)Pr(A∩B)​

The fraction of the world in which BBB happens that also has AAA in it. Everything else follows from this one definition.

Bayes in odds form

Pr⁡(H∣E)Pr⁡(Hˉ∣E)=Pr⁡(H)Pr⁡(Hˉ)×Pr⁡(E∣H)Pr⁡(E∣Hˉ)\frac{\Pr(H \mid E)}{\Pr(\bar H \mid E)} = \frac{\Pr(H)}{\Pr(\bar H)} \times \frac{\Pr(E \mid H)}{\Pr(E \mid \bar H)}Pr(Hˉ∣E)Pr(H∣E)​=Pr(Hˉ)Pr(H)​×Pr(E∣Hˉ)Pr(E∣H)​

Posterior odds are prior odds times the likelihood ratio. It is the fastest form to use at speed: no denominator to compute, and repeated independent evidence just multiplies more likelihood ratios in.

The law of total probability

Pr⁡(E)=∑iPr⁡(E∣Hi) Pr⁡(Hi)\Pr(E) = \sum_i \Pr(E \mid H_i)\,\Pr(H_i)Pr(E)=i∑​Pr(E∣Hi​)Pr(Hi​)

The denominator in Bayes, and often the whole question. Split the world into cases that do not overlap and cover everything, compute the probability of the evidence in each, and weight by how likely each case is.

Remember

  • Bayes follows from writing Pr⁡(A∩B)\Pr(A \cap B)Pr(A∩B) in both orders — derive it rather than recall it.

The distributions you have to know cold

Normal tail figures

Pr⁡(∣Z∣<1)≈68%,Pr⁡(∣Z∣<2)≈95%,Pr⁡(∣Z∣<3)≈99.7%\Pr(|Z| < 1) \approx 68\%, \quad \Pr(|Z| < 2) \approx 95\%, \quad \Pr(|Z| < 3) \approx 99.7\%Pr(∣Z∣<1)≈68%,Pr(∣Z∣<2)≈95%,Pr(∣Z∣<3)≈99.7%

And in the other direction, the z-scores for the intervals you are asked to quote:

Sums of exponentials

Tk=∑i=1kEi∼Gamma(k,λ),E[Tk]=kλ, Var⁡(Tk)=kλ2T_k = \sum_{i=1}^{k} E_i \sim \mathrm{Gamma}(k, \lambda), \quad \mathbb{E}[T_k] = \frac{k}{\lambda}, \ \operatorname{Var}(T_k) = \frac{k}{\lambda^2}Tk​=i=1∑k​Ei​∼Gamma(k,λ),E[Tk​]=λk​, Var(Tk​)=λ2k​

The waiting time until the kkk-th event of a Poisson process is a sum of kkk independent exponential gaps. It connects the three: Poisson counts, exponential gaps, gamma waiting times.

Remember

  • The table above, by recall — the questions test recognition, not derivation.

Linearity of expectation

The statement

E ⁣[∑i=1nXi]=∑i=1nE[Xi]\mathbb{E}\!\left[\sum_{i=1}^{n} X_i\right] = \sum_{i=1}^{n} \mathbb{E}[X_i]E[i=1∑n​Xi​]=i=1∑n​E[Xi​]

For any random variables X1,…,XnX_1,\dots,X_nX1​,…,Xn​ on the same probability space, and any constants aia_iai​:

With constants

E ⁣[∑iaiXi+b]=∑iai E[Xi]+b\mathbb{E}\!\left[\sum_i a_i X_i + b\right] = \sum_i a_i\,\mathbb{E}[X_i] + bE[i∑​ai​Xi​+b]=i∑​ai​E[Xi​]+b

Variance by indicators

Var⁡ ⁣(∑iIi)=∑iVar⁡(Ii)+2∑i<jCov⁡(Ii,Ij)\operatorname{Var}\!\left(\sum_i I_i\right) = \sum_i \operatorname{Var}(I_i) + 2\sum_{i<j}\operatorname{Cov}(I_i, I_j)Var(i∑​Ii​)=i∑​Var(Ii​)+2i<j∑​Cov(Ii​,Ij​)

Linearity gives the mean for free; the variance needs the pairwise covariances, which are usually just as easy because each needs only a joint probability of two events. For the hat-check problem it gives a variance of exactly one, for every n≥2n \ge 2n≥2.

Remember

  • Linearity of expectation holds for any random variables whatsoever, dependent or not.

Conditioning: the tower property

Law of total expectation

E[X]=E[ E[X∣Y] ]\mathbb{E}[X] = \mathbb{E}\big[\,\mathbb{E}[X \mid Y]\,\big]E[X]=E[E[X∣Y]]

The inner expectation is a function of YYY, so it is itself a random variable. The outer expectation averages it over YYY.

The form you will actually write

E[X]=∑yPr⁡(Y=y) E[X∣Y=y]\mathbb{E}[X] = \sum_{y} \Pr(Y = y)\,\mathbb{E}[X \mid Y = y]E[X]=y∑​Pr(Y=y)E[X∣Y=y]

A weighted average of the conditional answers, weighted by how likely each case is.

Law of total variance

Var⁡(X)=E[Var⁡(X∣Y)]⏟within-group+Var⁡(E[X∣Y])⏟between-group\operatorname{Var}(X) = \underbrace{\mathbb{E}\big[\operatorname{Var}(X \mid Y)\big]}_{\text{within-group}} + \underbrace{\operatorname{Var}\big(\mathbb{E}[X \mid Y]\big)}_{\text{between-group}}Var(X)=within-groupE[Var(X∣Y)]​​+between-groupVar(E[X∣Y])​​

Total variability splits into the average spread inside each scenario, plus the spread of the scenario averages.

Remember

  • E[X]=E[E[X∣Y]]\mathbb{E}[X] = \mathbb{E}[\mathbb{E}[X \mid Y]]E[X]=E[E[X∣Y]] for any YYY, so the only question is which YYY makes the inner expectation easy.

Recursive expected value and the re-roll family

Key rules

  • Condition on the first outcome; the rest is bookkeeping.
  • Keep an outcome when it beats the continuation value, and re-roll otherwise.
  • When the game repeats, the continuation value is the value of the game, which gives an equation in VVV alone.

Random walks and gambler’s ruin

The biased game

Pk=1−(q/p)k1−(q/p)NP_k = \frac{1 - (q/p)^{k}}{1 - (q/p)^{N}}Pk​=1−(q/p)N1−(q/p)k​

The probability of reaching NNN before 000, starting from kkk, when each step is up with probability ppp.

How long it lasts

E[steps]=k(N−k)when p=12\mathbb{E}[\text{steps}] = k(N - k) \quad \text{when } p = \tfrac{1}{2}E[steps]=k(N−k)when p=21​

The expected number of steps before absorption in a fair game — maximised in the middle, and larger than most people guess.

The martingale behind the biased formula

Mn=(qp)SnM_n = \left(\frac{q}{p}\right)^{S_n}Mn​=(pq​)Sn​

For a walk that steps up with probability ppp, (q/p)Sn(q/p)^{S_n}(q/p)Sn​ has constant expectation: one step multiplies it by p (q/p)+q (p/q)=1p\,(q/p) + q\,(p/q) = 1p(q/p)+q(p/q)=1. Optional stopping then gives the biased ruin probability in one line, just as SnS_nSn​ itself does for the fair game.

Remember

  • One equation per state, plus the boundaries, then solve.

Markov chains: states, transitions and hitting times

Turn arrows into a matrix

P=(0.80.20.50.5),(P2)ij=∑kPikPkjP=\begin{pmatrix}0.8&0.2\\0.5&0.5\end{pmatrix},\qquad (P^2)_{ij}=\sum_k P_{ik}P_{kj}P=(0.80.5​0.20.5​),(P2)ij​=k∑​Pik​Pkj​

Rows are today’s state and columns are tomorrow’s, in the order Clear, Storm. Matrix multiplication sums over the possible intermediate states. Every row of a transition matrix sums to one.

Remember

  • State what each node remembers. Test whether the future is independent of older history given that node.

Order statistics: maxima, minima and the gaps between

The two that factorise

Fmax⁡(x)=F(x)n,Fmin⁡(x)=1−(1−F(x))nF_{\max}(x) = F(x)^n, \qquad F_{\min}(x) = 1 - \big(1 - F(x)\big)^nFmax​(x)=F(x)n,Fmin​(x)=1−(1−F(x))n

The maximum is below xxx exactly when every draw is; the minimum is above xxx exactly when every draw is.

The full distribution of the $k$-th smallest

U(k)∼Beta(k, n+1−k),Var⁡(U(k))=k(n+1−k)(n+1)2(n+2)U_{(k)} \sim \mathrm{Beta}(k,\ n + 1 - k), \quad \operatorname{Var}(U_{(k)}) = \frac{k(n+1-k)}{(n+1)^2(n+2)}U(k)​∼Beta(k, n+1−k),Var(U(k)​)=(n+1)2(n+2)k(n+1−k)​

For nnn uniforms, the kkk-th smallest has a beta distribution. It gives variances as well as means, and shows that extremes are pinned down much more tightly than the median.

Remember

  • Fmax⁡=FnF_{\max} = F^nFmax​=Fn and Fmin⁡=1−(1−F)nF_{\min} = 1 - (1-F)^nFmin​=1−(1−F)n — always start from the CDF.

Simulation: making randomness you want out of randomness you have

The inverse transform

U∼Unif(0,1) ⇒ F−1(U)∼FU \sim \text{Unif}(0,1) \ \Rightarrow \ F^{-1}(U) \sim FU∼Unif(0,1) ⇒ F−1(U)∼F

Feed a uniform through the inverse CDF and you get the distribution you wanted. This is the general answer to "how would you sample from this?", and it is why a uniform generator is all a library needs.

Box–Muller

Z1=−2ln⁡U1cos⁡(2πU2),Z2=−2ln⁡U1sin⁡(2πU2)Z_1 = \sqrt{-2\ln U_1}\cos(2\pi U_2), \quad Z_2 = \sqrt{-2\ln U_1}\sin(2\pi U_2)Z1​=−2lnU1​​cos(2πU2​),Z2​=−2lnU1​​sin(2πU2​)

Two independent uniforms give two independent standard normals. It works because the joint normal density is rotationally symmetric: the angle is uniform and the squared radius is exponential with mean two, and both are easy to generate.

Control variates

θ^cv=Yˉ−c (Xˉ−μX),c∗=Cov⁡(X,Y)Var⁡(X)\hat\theta_{\text{cv}} = \bar Y - c\,(\bar X - \mu_X), \quad c^* = \frac{\operatorname{Cov}(X, Y)}{\operatorname{Var}(X)}θ^cv​=Yˉ−c(Xˉ−μX​),c∗=Var(X)Cov(X,Y)​

If you can simulate a quantity XXX whose mean you know exactly and which moves with the thing you want, subtract its error. With the optimal coefficient the variance falls by the factor 1−ρ21 - \rho^21−ρ2 — a correlation of 0.90.90.9 cuts it by 81%81\%81%, worth five times as many paths.

Remember

  • F−1(U)F^{-1}(U)F−1(U) samples any distribution; −ln⁡U/λ-\ln U/\lambda−lnU/λ is the exponential case.

Detailed formula cards

  • Linearity of expectation
  • Indicator expectation
  • Law of total expectation
  • Law of total variance
  • Second moment identity
  • Value of a repeatable game with a fee

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.