arXiv:2510.15513cs.CL2025-10EMNLP

测试大模型对时间参照的一致性,发现其更倾向序列时间而非绝对时间。

Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?

  • 构建新基准TEMP-ReCon,评估模型在多语言时间推理中的一致性
  • 实测显示大模型在绝对时间参照上一致性不足,序列时间更受偏好
  • 提出新方法UnTRaP,通过路径对齐提升时间参照一致性

大型语言模型(LLMs)在法律、医疗、金融等时间敏感领域的广泛应用,要求其不仅事实准确,还需在时间维度上保持一致。然而,当前对大模型时间一致性的研究仍十分匮乏。本文提出新的基准TEMP-ReCon,涵盖英语、法语和罗马尼亚语等多种语言环境,用于评估开源与闭源模型的时间参照一致性。实验结果表明,大模型在绝对时间引用上的表现较差,更倾向于依赖序列时间关系。为此,我们提出基于推理路径对齐的UnTRaP模型,有效提升了时间参照一致性。实验证明,该方法优于多个基线模型。

原文摘要 · Abstract (English)

The increasing acceptance of large language models (LLMs) as an alternative to knowledge sources marks a significant paradigm shift across various domains, including time-sensitive fields such as law, healthcare, and finance. To fulfill this expanded role, LLMs must not only be factually accurate but also demonstrate consistency across temporal dimensions, necessitating robust temporal reasoning capabilities. Despite this critical requirement, efforts to ensure temporal consistency in LLMs remain scarce including noticeable absence of endeavors aimed at evaluating or augmenting LLMs across temporal references in time-sensitive inquiries. In this paper, we seek to address this gap by introducing a novel benchmark entitled temporal referential consistency, accompanied by a resource TEMP-ReCon designed to benchmark a wide range of both open-source and closed-source LLMs with various linguistic contexts characterized by differing resource richness (including English, French, and Romanian). The findings emphasis that LLMs do exhibit insufficient temporal referent consistency. To address this, we propose \newmodel, a reasoning path alignment-based model that aims to enhance the temporal referential consistency of LLMs. Our empirical experiments substantiate the efficacy of UnTRaP compared to several baseline models.

大模型时间推理一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。