arXiv:2501.03040cs.LGcs.CL2025-01ACL被引 26

构建新基准测试大模型时间推理能力,发现其表现不佳且依赖记忆。

ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events

  • 设计16项任务,基于艾伦区间关系评估事件时间顺序与算术。
  • 7个主流模型在抽象与真实数据上表现均不理想,平均准确率不足50%。
  • 适合关注时序理解、模型可解释性的研究者使用。

大型语言模型(LLMs)在多种自然语言任务中取得显著进展,但在推理与算术方面仍面临挑战。时间推理作为自然语言理解的关键部分,日益受到关注。然而,对艾伦区间关系(如先于、后于、包含)——时间关系的基础框架——的全面评估仍显不足。为此,我们提出ChronoSense,一个用于评估LLMs时间理解的新基准。该基准包含16项任务,聚焦于识别两个时间事件之间的艾伦关系及时间算术,涵盖抽象事件与来自Wikidata的真实世界数据。我们使用该基准评估了7个近期的LLM,结果表明模型对艾伦关系,甚至对称关系的处理方式差异显著。此外,研究发现模型可能依赖记忆回答时间相关问题。总体而言,模型表现低下凸显了提升其时间理解能力的必要性,而ChronoSense为未来研究提供了可靠框架。数据集与源代码已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success in various NLP tasks, yet they still face significant challenges in reasoning and arithmetic. Temporal reasoning, a critical component of natural language understanding, has raised increasing research attention. However, comprehensive testing of Allen's interval relations (e.g., before, after, during) -- a fundamental framework for temporal relationships -- remains underexplored. To fill this gap, we present ChronoSense, a new benchmark for evaluating LLMs' temporal understanding. It includes 16 tasks, focusing on identifying the Allen relation between two temporal events and temporal arithmetic, using both abstract events and real-world data from Wikidata. We assess the performance of seven recent LLMs using this benchmark and the results indicate that models handle Allen relations, even symmetrical ones, quite differently. Moreover, the findings suggest that the models may rely on memorization to answer time-related questions. Overall, the models' low performance highlights the need for improved temporal understanding in LLMs and ChronoSense offers a robust framework for future research in this area. Our dataset and the source code are available at https://github.com/duyguislakoglu/chronosense.

时间推理语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。