arXiv:2505.19533cs.LG2025-05Conference of the …被引 5

测试大模型在时间限制下的推理能力,发现它们常‘预知’未来。

ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

  • 设计新评测任务,强制模型按时间线推理
  • 多数模型仍受未来信息干扰,泄漏率高
  • 适合研究时序推理与可信AI的开发者

大语言模型在事前推理中面临挑战:需在无法获取未来信息的情况下进行分析或预测。即使通过显式提示设定时间截止,模型仍会受到内部存储的未来事件知识影响。本文提出一项新任务与基准,用于评估模型在时间约束下的推理能力,涵盖股票预测、维基百科事件预测、科学论文预测和问答任务,旨在检验模型在时间截断条件下的事实性知识。采用泄漏率量化模型对截止时间后信息的依赖程度。实验表明,当前主流提示策略下,模型在各类任务中均难以一致遵守时间限制,暴露出事前推理的持续难题。该基准为提升大模型时序推理能力提供了评估框架,有助于推动其在时间敏感场景中的应用发展。

原文摘要 · Abstract (English)

Large language models (LLMs) face significant challenges in ex-ante reasoning, where analysis, inference, or predictions must be made without access to information from future events. Even with explicit prompts enforcing temporal cutoffs, LLMs often generate outputs influenced by internalized knowledge of events beyond the specified cutoff. This paper introduces a novel task and benchmark designed to evaluate the ability of LLMs to reason while adhering to such temporal constraints. The benchmark includes a variety of tasks: stock prediction, Wikipedia event prediction, scientific publication prediction, and Question Answering (QA), designed to assess factual knowledge under temporal cutoff constraints. We use leakage rate to quantify models' reliance on future information beyond cutoff timestamps. Experimental results reveal that LLMs struggle to consistently adhere to temporal cutoffs across common prompting strategies and tasks, demonstrating persistent challenges in ex-ante reasoning. This benchmark provides a potential evaluation framework to advance the development of LLMs' temporal reasoning ability for time-sensitive applications.

大模型推理评测时间约束可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。