过参数化模型在数据分布变化时预测能力下降,金融建模需警惕。
Overparametrized models with posterior drift
- 研究过参数模型在训练与测试分布不同时的后验漂移问题。
- 小带宽模型收益差异大,大带宽虽稳定但风险调整后回报差。
- 适合关注长期投资与模型稳健性的金融量化研究者。
本文研究过参数化机器学习模型中后验漂移对样本外预测精度的影响。我们发现当数据生成过程的载荷在训练与测试阶段发生变化时,模型性能显著下降。这一问题在可能发生制度变迁的场景中尤为关键,如金融市场。应用于股票溢价预测,结果表明市场择时策略对子周期和控制模型复杂度的带宽参数极为敏感。对普通投资者而言,聚焦15年持有期时,小带宽模型产生极不稳定的收益,而大带宽模型虽更一致,但从风险调整回报角度看吸引力不足。总体而言,我们的发现提示在股票市场预测中应谨慎使用大型线性模型。
原文摘要 · Abstract (English)
This paper investigates the impact of posterior drift on out-of-sample forecasting accuracy in overparametrized machine learning models. We document the loss in performance when the loadings of the data generating process change between the training and testing samples. This matters crucially in settings in which regime changes are likely to occur, for instance, in financial markets. Applied to equity premium forecasting, our results underline the sensitivity of a market timing strategy to sub-periods and to the bandwidth parameters that control the complexity of the model. For the average investor, we find that focusing on holding periods of 15 years can generate very heterogeneous returns, especially for small bandwidths. Large bandwidths yield much more consistent outcomes, but are far less appealing from a risk-adjusted return standpoint. All in all, our findings tend to recommend cautiousness when resorting to large linear models for stock market predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。