arXiv:2506.03780q-fin.STcs.LG2025-06被引 3

揭示金融高维学习的理论极限,发现真实成功源于低复杂度信号。

High-Dimensional Learning in Finance

  • 发现随机傅里叶特征中样本内标准化会改变核函数本质
  • 证明在典型参数下需25-30年数据才可能突破学习下限
  • 适合关注金融预测可信性与模型复杂度的从业者

机器学习在金融预测中使用大规模过参数化模型展现出良好前景。本文为理解此类方法何时何地取得成功提供理论基础与实证验证。研究高维学习在金融中的两个关键方面:首先,证明随机傅里叶特征实现中的样本内标准化从根本上改变了高斯核逼近,将平移不变核替换为依赖训练集的替代形式;其次,建立信息论下界,揭示无论估计器多么复杂,何时可靠学习根本不可能。对多项式下界的定量校准显示,采用典型参数(如12,000个特征、12个月观测、决定系数R² 2-3%),要摆脱该下界所需样本量超过25-30年数据——远超实际滚动窗口使用范围。因此,观测到的样本外成功必源于低复杂度伪象,而非预期的高维机制。

原文摘要 · Abstract (English)

Recent advances in machine learning have shown promising results for financial prediction using large, over-parameterized models. This paper provides theoretical foundations and empirical validation for understanding when and how these methods achieve predictive success. I examine two key aspects of high-dimensional learning in finance. First, I prove that within-sample standardization in Random Fourier Features implementations fundamentally alters the underlying Gaussian kernel approximation, replacing shift-invariant kernels with training-set dependent alternatives. Second, I establish information-theoretic lower bounds that identify when reliable learning is impossible no matter how sophisticated the estimator. A detailed quantitative calibration of the polynomial lower bound shows that with typical parameter choices, e.g., 12,000 features, 12 monthly observations, and R-square 2-3%, the required sample size to escape the bound exceeds 25-30 years of data--well beyond any rolling-window actually used. Thus, observed out-of-sample success must originate from lower-complexity artefacts rather than from the intended high-dimensional mechanism.

金融预测高维学习模型复杂度理论边界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。