ExpertMultiple choice
One step of gradient descent · Part 4 of 4
A logistic model predicts , where , with one weight and no bias. It starts at and trains on a single example with and label , using the cross-entropy loss .
You keep training on this one example, or on any data set the model separates perfectly, with no regularisation. What happens to ?
- AIt converges to the value where the loss is exactly zero
- BIt grows without bound as the loss creeps towards zero
- CIt oscillates, because the gradient changes sign when p passes 0.5
- DIt stops at w = 0.5 because the loss is convex
The worked solution is in Premium
The answer, the full working and the one idea to take away – for this and all 1,322 questions in the bank. Answer it in practice and your working is marked, with a known mistake named when you make one.
Learn the method
More machine learning questions
- Gradient descent with momentum uses a learning rate of 0.01 and β = 0.9.Applied
- Gradient descent on f(x) = x² starts at x = 1 with learning rate 0.1.Applied
- A quadratic loss has Hessian eigenvalues 4 and 0.5.Applied
- You increase the batch size from 32 to 512 without changing anything else. What happens?Advanced
- Adam’s first-moment estimate starts at 0 with β₁ = 0.9.Advanced
- You deepen your trees from depth 3 to depth 12. What happens to bias and variance?Foundation