测试大模型跨历法时间推理能力,发现准确率仅34.5%。
SPAN: Benchmarking and Improving Cross-Calendar Temporal Reasoning of Large Language Models
- 构建动态生成的跨历法推理基准,支持多历法、双类型、双格式题目
- 模型平均准确率34.5%,最高未超80%,暴露时间推理短板
- 提出工具增强的时序智能体,准确率达95.31%,适合长期时间任务研究者
我们提出SPAN,一个跨历法时间推理基准,要求大语言模型完成历法内推理与历法间转换。SPAN包含十种跨历法推理方向、两种推理类型和两种问题格式,覆盖六种历法。为实现时间敏感且无污染评估,我们设计基于模板的动态实例生成协议,可在用户指定的格里高利历日期上进行测试。我们在1960至2060年间跨多个日期对开源与闭源最先进模型进行了广泛实验。结果显示,这些模型平均准确率仅为34.5%,且无一超过80%,表明该任务仍具挑战性。通过深入分析推理类型、问题格式与推理方向,我们识别出两大障碍:未来日期退化与历法不对称偏见。为进一步提升跨历法时间推理能力,我们开发了基于大模型的时序智能体(Time Agent),利用工具增强代码生成。实证结果表明,Time Agent平均准确率达95.31%,优于多个基线模型,凸显工具增强代码生成在推进跨历法时间推理方面的潜力。我们希望本工作能推动更具备时间与文化适应性的大模型发展。
原文摘要 · Abstract (English)
We introduce SPAN, a cross-calendar temporal reasoning benchmark, which requires LLMs to perform intra-calendar temporal reasoning and inter-calendar temporal conversion. SPAN features ten cross-calendar temporal reasoning directions, two reasoning types, and two question formats across six calendars. To enable time-variant and contamination-free evaluation, we propose a template-driven protocol for dynamic instance generation that enables assessment on a user-specified Gregorian date. We conduct extensive experiments on both open- and closed-source state-of-the-art (SOTA) LLMs over a range of dates spanning 100 years from 1960 to 2060. Our evaluations show that these LLMs achieve an average accuracy of only 34.5%, with none exceeding 80%, indicating that this task remains challenging. Through in-depth analysis of reasoning types, question formats, and temporal reasoning directions, we identify two key obstacles for LLMs: Future-Date Degradation and Calendar Asymmetry Bias. To strengthen LLMs' cross-calendar temporal reasoning capability, we further develop an LLM-powered Time Agent that leverages tool-augmented code generation. Empirical results show that Time Agent achieves an average accuracy of 95.31%, outperforming several competitive baselines, highlighting the potential of tool-augmented code generation to advance cross-calendar temporal reasoning. We hope this work will inspire further efforts toward more temporally and culturally adaptive LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。