arXiv:2608.23761stat.MLcs.LG2026-08

发现股票收益预测中过度拟合未必有益,模型最终仍不如简单历史平均。

(Mis)Understanding Benign Overfitting in Equity Return Prediction

论文配图:(Mis)Understanding Benign Overfitting in Equity Return Prediction
图 1 · 摘自论文原文
  • 用无正则化与有正则化模型验证双下降现象
  • 参数量远超样本量时性能差距消失,但均不及历史均值
  • 揭示机器学习在无真实信号时会退化为简单基准

高度过参数化的模型在复杂领域中常能良好预测,尽管它们插值训练数据,这挑战了经典的偏差-方差权衡。我们研究该‘良性过拟合’现象是否适用于股票收益预测。与近期统计理论一致,我们观察到两个关键现象:第一,无正则化模型的预测风险呈现双下降模式;第二,虽然最优带正则化模型始终优于无正则化版本,但在参数量与观测数之比极大时,二者性能差距趋于消失。然而,两者均无法超越简单的历史平均表现。这一实证结果与零斜率系数假设下的渐近分析一致,表明标准股票预测因子本质上缺乏真实预测能力——即使在高度灵活的非线性机器学习架构中也是如此。这些发现调和了现代与传统机器学习在资产定价中的矛盾:当不存在真实信号时,两者最终都退化为历史平均基准。

原文摘要 · Abstract (English)

Highly overparameterized models often predict well despite interpolating training data in complex domains, challenging the classical bias--variance tradeoff. We investigate whether this ``benign overfitting'' phenomenon extends to equity return prediction. Consistent with recent statistical theory, we document two key phenomena: first, a double descent pattern in the ridgeless model's prediction risk; and second, that while the optimal ridge model consistently outperforms its ridgeless counterpart, this performance gap becomes negligible at large parameter-to-observation ratios. Ultimately, however, both models fail to outperform a simple historical average. This empirical evidence aligns with our asymptotic results under the null hypothesis of zero slope coefficients, suggesting that standard equity predictors lack true forecasting power---even within highly flexible, nonlinear machine learning architectures. These findings reconcile modern and classical machine learning in asset pricing: in the absence of a true signal, they asymptotically collapse to the historical average benchmark.

股票预测过拟合双下降机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。