arXiv:2601.07148cs.CLcs.AI2026-01被引 1

用时间谜题测试大模型迭代推理能力,发现其工具使用仍不靠谱。

Measuring Iterative Temporal Reasoning with Time Puzzles

  • 设计可自动生成的时间谜题,结合日历关系与事实锚点
  • 最佳模型无工具时仅55.3%准确率,依赖搜索仍难达标
  • 显式日期约束能大幅提升性能,揭示工具调用缺陷

工具使用(如网络搜索)已成为主流大语言模型的标准能力。然而,现有评估基准主要在静态、不使用工具的环境下进行时间推理测试,无法真实反映大模型在实际中的表现。本文提出Time Puzzles,一种基于约束的时间推断任务,用于评估大模型在使用工具时的迭代时间推理能力。每个谜题结合事实性时间锚点与跨文化日历关系,可能有唯一或多个有效日期。谜题通过算法生成,支持可控且持续的评估。在13个大模型中,即使最优模型GPT-5在不使用工具的情况下也仅达55.3%准确率,尽管所涉信息均易检索。虽然网络搜索能提升表现,但当约束被重写为显式日期时,模型性能显著提升,无需事实查询。结果揭示了大模型在可靠工具使用方面存在明显短板。

原文摘要 · Abstract (English)

Tool use, such as web search, has become a standard capability even in freely available large language models (LLMs). However, existing benchmarks evaluate temporal reasoning mainly in static, non-tool-using settings, which poorly reflect how LLMs perform temporal reasoning in practice. We introduce Time Puzzles, a constraint-based date inference task for evaluating iterative temporal reasoning with tools. Each puzzle combines factual temporal anchors with (cross-cultural) calendar relations and may admit one or multiple valid dates. The puzzles are algorithmically generated, enabling controlled and continual evaluation. Across 13 LLMs, even the best model (GPT-5) achieves only 55.3% accuracy without tools, despite using easily searchable facts. While web search improves performance, models perform substantially better when constraints are rewritten with explicit dates, removing the need for factual lookup. These results reveal a gap in reliable tool use for iterative temporal reasoning.

时间推理工具使用评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。