AppliedNumeric answer
Gradient descent with momentum uses a learning rate of and . What is the effective step size on a constant gradient?
Answer with a number. Fractions, powers and expressions like 23/6 or C(52,5) are read correctly in practice.
The worked solution is in Premium
The answer, the full working and the one idea to take away – for this and all 1,322 questions in the bank. Answer it in practice and your working is marked, with a known mistake named when you make one.
Learn the method
Reported in interviews at
Jane Street, Citadel, Hudson River Trading, Jump Trading, D. E. Shaw
More machine learning questions
- One step of gradient descent, part 1 of 4Foundation
- Gradient descent on f(x) = x² starts at x = 1 with learning rate 0.1.Applied
- A quadratic loss has Hessian eigenvalues 4 and 0.5.Applied
- One step of gradient descent, part 2 of 4Applied
- You increase the batch size from 32 to 512 without changing anything else. What happens?Advanced
- Adam’s first-moment estimate starts at 0 with β₁ = 0.9.Advanced