arXiv:2504.05258cs.LGcs.AI2025-04ACL被引 11

通过时间线自省提升大模型的时间推理能力,让机器更懂事件先后。

Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models

  • 构建时间线并迭代反思,增强时间关系理解
  • 小模型经训练后超越大闭源模型表现
  • 适合需要精准时间逻辑的问答与历史分析场景

大语言模型在生成连贯文本、理解上下文和执行推理任务方面表现出色,但在时间推理上仍存在短板,难以处理事件排序、持续时间及时间间关系等信息。这些能力对问答、日程安排和历史分析等应用至关重要。本文提出TISER框架,通过多阶段流程结合时间线构建与迭代自省,提升大模型的时间推理能力。该方法利用测试时扩展(test-time scaling)延长推理路径,使模型更有效地捕捉复杂时间依赖。实验表明,TISER在多个基准测试中达到领先水平,包括分布外数据集,并发现小型开源模型经训练后可超越大型闭源模型在挑战性时间推理任务上的表现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as powerful tools for generating coherent text, understanding context, and performing reasoning tasks. However, they struggle with temporal reasoning, which requires processing time-related information such as event sequencing, durations, and inter-temporal relationships. These capabilities are critical for applications including question answering, scheduling, and historical analysis. In this paper, we introduce TISER, a novel framework that enhances the temporal reasoning abilities of LLMs through a multi-stage process that combines timeline construction with iterative self-reflection. Our approach leverages test-time scaling to extend the length of reasoning traces, enabling models to capture complex temporal dependencies more effectively. This strategy not only boosts reasoning accuracy but also improves the traceability of the inference process. Experimental results demonstrate state-of-the-art performance across multiple benchmarks, including out-of-distribution test sets, and reveal that TISER enables smaller open-source models to surpass larger closed-weight models on challenging temporal reasoning tasks.

时间推理自省机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。