arXiv:2505.12891cs.AIcs.CL2025-05NeurIPS被引 14

构建多层级时序推理基准,评估大模型在真实场景中的时间理解能力。

TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios

  • 设计三级任务体系,覆盖真实世界的时序信息、事件动态与社交依赖挑战。
  • 包含38,522个问答对,分三子数据集:TIME-Wiki、TIME-News、TIME-Dial。
  • 提供轻量版人工标注数据集,助力未来标准化评估与研究。

时序推理对大语言模型理解现实世界至关重要。现有研究忽视了真实场景中的三大挑战:(1)密集的时序信息,(2)快速变化的事件动态,(3)社会互动中的复杂时序依赖。为此,我们提出多层级基准TIME,用于评估大模型在真实场景中的时序推理能力。TIME包含38,522个问答对,涵盖3个层级与11个细粒度子任务,包含三个反映不同现实挑战的子数据集:TIME-Wiki、TIME-News和TIME-Dial。我们在多种推理与非推理模型上进行了广泛实验,深入分析了不同真实场景与任务下的时序推理表现,并总结了测试时缩放对时序推理能力的影响。此外,我们发布了TIME-Lite,一个由人工标注的子集,以推动未来研究与标准化评估。代码已开源至https://github.com/sylvain-wei/TIME,数据集可在https://huggingface.co/datasets/SylvainWei/TIME获取,项目主页为https://sylvain-wei.github.io/TIME/。

原文摘要 · Abstract (English)

Temporal reasoning is pivotal for Large Language Models (LLMs) to comprehend the real world. However, existing works neglect the real-world challenges for temporal reasoning: (1) intensive temporal information, (2) fast-changing event dynamics, and (3) complex temporal dependencies in social interactions. To bridge this gap, we propose a multi-level benchmark TIME, designed for temporal reasoning in real-world scenarios. TIME consists of 38,522 QA pairs, covering 3 levels with 11 fine-grained sub-tasks. This benchmark encompasses 3 sub-datasets reflecting different real-world challenges: TIME-Wiki, TIME-News, and TIME-Dial. We conduct extensive experiments on reasoning models and non-reasoning models. And we conducted an in-depth analysis of temporal reasoning performance across diverse real-world scenarios and tasks, and summarized the impact of test-time scaling on temporal reasoning capabilities. Additionally, we release TIME-Lite, a human-annotated subset to foster future research and standardized evaluation in temporal reasoning. The code is available at https://github.com/sylvain-wei/TIME , the dataset is available at https://huggingface.co/datasets/SylvainWei/TIME , and the project page link is https://sylvain-wei.github.io/TIME/ .

时序推理大模型评估多层级基准真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。