arXiv:2512.12735stat.MLcs.LG2025-12

机器学习在金融数据中预测能力被低估,因样本有限导致固有误差。

Limits To (Machine) Learning

  • 提出学习上限差距(LLG),量化模型在有限样本下的性能天花板。
  • 实证显示金融变量的LLG值大,说明传统ML低估了真实可预测性。
  • 适用于研究金融预测、资产定价与模型过拟合问题的学者。

机器学习方法虽灵活,但其对真实数据生成过程的逼近能力受样本量限制。本文定义了一个普适下界——学习上限差距(LLG),量化模型经验拟合与总体基准之间的不可避免差距。因此,要恢复真实的总体$R^2$,需对观测到的预测性能进行该边界的修正。基于包括超额收益、收益率、信用利差和估值比率在内的广泛变量,我们发现隐含的LLG值较大,表明标准机器学习方法在金融数据中显著低估了真实可预测性。此外,本文推导出基于LLG的经典Hansen-Jagannathan(1991)边界改进形式,分析了其在一般均衡设定下参数学习的影响,并证明LLG可自然产生过度波动。

原文摘要 · Abstract (English)

Machine learning (ML) methods are highly flexible, but their ability to approximate the true data-generating process is fundamentally constrained by finite samples. We characterize a universal lower bound, the Limits-to-Learning Gap (LLG), quantifying the unavoidable discrepancy between a model's empirical fit and the population benchmark. Recovering the true population $R^2$, therefore, requires correcting observed predictive performance by this bound. Using a broad set of variables, including excess returns, yields, credit spreads, and valuation ratios, we find that the implied LLGs are large. This indicates that standard ML approaches can substantially understate true predictability in financial data. We also derive LLG-based refinements to the classic Hansen and Jagannathan (1991) bounds, analyze implications for parameter learning in general-equilibrium settings, and show that the LLG provides a natural mechanism for generating excess volatility.

机器学习金融预测模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。