AdvancedMultiple choice
Why did rectified linear units largely replace sigmoids in deep networks?
- AA sigmoid saturates, so its gradient vanishes and deep stacks stop learning
- BRectifiers are smoother, which helps the optimiser
- CRectifiers bound the activations, which stabilises training
- DRectifiers guarantee convexity of the loss
The worked solution is in Premium
The answer, the full working and the one idea to take away – for this and all 1,322 questions in the bank. Answer it in practice and your working is marked, with a known mistake named when you make one.
Reported in interviews at
More machine learning questions
- How many learnable parameters does a fully connected layer from 100 inputs to…Foundation
- A softmax output layer receives logits (2,1,0).Foundation
- A batch normalisation layer follows a fully connected layer of 128 units.Applied
- A convolutional layer uses 3 × 3 kernels, 16 input channels and 32 output…Applied
- You deepen your trees from depth 3 to depth 12. What happens to bias and variance?Foundation
- A node holds 40 positive and 60 negative examples. What is its Gini impurity?Foundation