arXiv:2605.30393cs.LGcs.AI2026-05中稿 · ICML被引 1

发现大模型会记忆公开金融数据,导致评测结果失真。

NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models

论文配图:NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models
图 1 · 摘自论文原文
  • 用接口探测与白盒验证结合,检测模型对公开数值的过拟合。
  • 顶级模型对金融因子相关性达0.97-0.99,但新数据解析率仅21-57%。
  • 提示词防御可阻断99.8%攻击,且几乎无成本,适合实际部署。

公开的数值基准在预训练中出现,因此基于时间的评估可能测量的是记忆召回而非泛化能力。我们提出NumLeak,一种结合API边界探测与开放因果语言模型白盒验证的测量框架。顶级前沿LLM在3个随机种子下对Fama-French市场超额收益的皮尔逊相关系数达到0.97-0.99,对五个同类因子的误差在±25bps内;美国失业率、CPI通胀和NOAA温度数据也表现出类似保真度。在近期发布数据集上,解析率降至21-57%,但回答月份的相关性仍维持约0.99,拒绝-回忆不对称性表明存在记忆通道。白盒实验重现剂量反应关系,对数概率排序能检测出开放生成忽略的记忆行为,说明封闭接口探测低估了该通道。一个针对Sonnet的“日期到市场情绪”回归模型在真实市场超额收益上的相关系数为0.74,但在剔除模型自身记忆后骤降至0.02。一条简短的系统提示即可阻止99.8%的非自适应单轮后缀攻击,在概念性和历史叙事任务上几乎零成本。

原文摘要 · Abstract (English)

Public numeric benchmarks appear in pretraining, so an evaluation that conditions on a date may be measuring memorized recall rather than out-of-sample skill. We introduce NumLeak, a measurement framework that combines API-boundary probes on production models with a white-box controlled validation on an open causal LM. Top-tier frontier LLMs recall the Fama-French market excess return at 3-seed pooled Pearson r=0.97-0.99 while staying within 0.15 within-25bps on the five sibling factors; comparable fidelity appears on U.S. unemployment, CPI inflation, and NOAA temperature. On a recent-release holdout, parse rate collapses to 21-57% but r stays at approximately 0.99 on months answered, the refuse-or-recall asymmetry a memorized channel predicts. The white-box experiment reproduces the dose-response, and logprob ranking detects memorization that open-ended generation misses, implying closed-API black-box probes understate the channel. A Sonnet "date to market-sentiment" regression that correlates with true Mkt-RF at r=0.74 collapses to r=0.02 once the model's own recall is residualized out. A one-line system-prompt defense blocks 99.8% of a non-adaptive single-turn suffix attack set at near-zero utility cost on conceptual and historical-narrative queries

大模型安全记忆泄露金融预测评测偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。