arXiv:2601.13717cs.CLcs.AI2026-01被引 5

Simulated Ignorance无法真实模拟模型无知,导致预测能力评估失效。

Simulated Ignorance Fails: A Systematic Study of LLM Behaviors on Forecasting Problems Before Model Knowledge Cutoff

  • 用提示词让模型假装不知情,但实际仍受旧知识干扰。
  • 仿真无知与真实无知性能差距达52%,且推理越优越难抑制旧知。
  • 现有回溯评估方法存在根本缺陷,不适合测试模型预测能力。

评估大模型的预测能力面临根本矛盾:前瞻性评估虽严谨但延迟高;回溯预测(RF)——在已解决事件上评估——则因模型知识截止时间越来越新而可用数据迅速减少。模拟无知(SI)通过指令要求模型屏蔽截止前的知识,被视为潜在解决方案。本文首次系统检验了SI能否逼近真实无知(TI)。在477个竞赛级问题和9个模型上发现,SI系统性失败:(1)截止指令导致SI与TI之间存在52%的性能差距;(2)链式思维推理无法有效抑制先验知识,即使推理过程未显式提及截止后信息;(3)推理优化模型反而在SI中表现更差,尽管其推理质量更高。结果表明,仅靠提示词无法可靠‘回退’模型知识。我们得出结论:基于预截止事件的回溯评估方法论存在缺陷,应避免使用基于SI的回溯设置来评测预测能力。

原文摘要 · Abstract (English)

Evaluating LLM forecasting capabilities is constrained by a fundamental tension: prospective evaluation offers methodological rigor but prohibitive latency, while retrospective forecasting (RF) -- evaluating on already-resolved events -- faces rapidly shrinking clean evaluation data as SOTA models possess increasingly recent knowledge cutoffs. Simulated Ignorance (SI), prompting models to suppress pre-cutoff knowledge, has emerged as a potential solution. We provide the first systematic test of whether SI can approximate True Ignorance (TI). Across 477 competition-level questions and 9 models, we find that SI fails systematically: (1) cutoff instructions leave a 52% performance gap between SI and TI; (2) chain-of-thought reasoning fails to suppress prior knowledge, even when reasoning traces contain no explicit post-cutoff references; (3) reasoning-optimized models exhibit worse SI fidelity despite superior reasoning trace quality. These findings demonstrate that prompts cannot reliably "rewind" model knowledge. We conclude that RF on pre-cutoff events is methodologically flawed; we recommend against using SI-based retrospective setups to benchmark forecasting capabilities.

大模型评估预测能力知识截止

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。