FoundationMultiple choice
Cross-validating a return forecaster · Part 1 of 3
You train a model on ten years of daily data. Each feature uses a trailing 20-day window, and each label is the stock’s return over the next 5 days. You plan to choose hyperparameters by cross-validation.
Why is ordinary shuffled k-fold cross-validation misleading here?
- AOverlapping labels and features let test days leak into training
- BTen folds are too few for daily data
- CShuffling changes the class balance between folds
- Dk-fold always underestimates performance because each model sees less data
The worked solution is in Premium
The answer, the full working and the one idea to take away – for this and all 1,322 questions in the bank. Answer it in practice and your working is marked, with a known mistake named when you make one.
Learn the method
More machine learning questions
- A classifier has 30 true positives, 20 false positives and 10 false negatives.Foundation
- A classifier predicts probability 0.8 for an example that is in fact positive.Foundation
- Why is ordinary shuffled k-fold cross-validation wrong on a financial time series?Applied
- A classifier assigns scores at random. What is its expected AUC?Applied
- A classifier flags 500 cases, of which 40 are true positives, and misses 60 positives.Applied
- An anomaly detector has sensitivity 90% and specificity 95%.Applied