arXiv:2605.24564cs.AIcs.CE2026-05被引 2

用大模型回测金融数据会因记忆历史而失真,新方法可消除这种偏差。

Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models

论文配图:Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models
图 1 · 摘自论文原文
  • 通过对抗性提示发现模型对历史事件的记忆,动态调整推理强度。
  • 最大模型回测收益修正达-67.1%,外样本表现仍稳定在基线附近。
  • 适合金融量化研究者,让大模型回测结果更可信。

在历史金融数据上回测大型语言模型(LLMs)时,若其预训练数据包含评估事件,则结果不可靠。例如2024年训练的模型可能已记住2018–2020年的股价走势。我们称此为参数化前瞻偏差,并提出FinCAD——一种无需重训练的推理时适应方法,基于上下文感知解码(CAD)来削弱对记忆历史结果的依赖。FinCAD结合对抗性偏差探测管道,学习特定模型的记忆激活提示,并引入实体与日期自适应规则,依据每(实体, 日期)的置信度信号调节CAD强度。在五个7–14B规模的LLM和五个大盘股上测试,最大模型级样本内收益修正达-67.1%。对三个较大模型,2025年外样本收益保持在±$8K内,平均夏普比率波动不超过±0.10;四款模型的通用基准准确率仍为正或低于-1.7分。在十一模型排行榜中,FinCAD将子集平均样本内/外样本斯皮尔曼相关性从+0.779提升至+0.846,使排名更贴近截断后真实表现。

原文摘要 · Abstract (English)

Backtesting large language models (LLMs) on historical financial data is unreliable when their pre-training data include the evaluated events. An LLM trained in 2024 may already encode how stocks moved during 2018-2020. We name this failure parametric look-ahead bias and propose FinCAD, an inference-time adaptation of Context-Aware Decoding that attenuates contributions from memorised historical outcomes without retraining. FinCAD pairs an adversarial bias-discovery pipeline that learns a model-specific memory-activating prior prompt with an entity- and date-adaptive rule that scales the CAD strength using a per-(entity, date) confidence signal. Across five 7-14B LLMs and five mega-cap equities, the largest model-level mean in-sample return correction is -67.1%. For the three larger models, 2025 out-of-sample returns remain within $8K and mean Sharpe within $\pm$0.10 of baseline; mean general-benchmark accuracy remains positive or within -1.7 points for four of five models. On an eleven-model leaderboard, FinCAD raises the subset-averaged in-sample/out-of-sample Spearman correlation from +0.779 to +0.846, yielding rankings that are more closely aligned with post-cutoff performance.

金融回测大模型偏差LLM应用量化研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。