测试大模型对时间相关事实的回忆能力,发现越复杂的模型反而越容易出错。
Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time
- 构建覆盖2018至2024年、超8000条事件的时序数据集,标注精确到天。
- 基础模型在时间敏感问题上表现优于指令微调和合成训练模型。
- 模型对改写后的事实仍易出错,暴露时间一致性关键缺陷。
谁是美国总统?答案随提问时间而变化。尽管大语言模型(LLMs)在多种推理任务中被评估,但往往忽略了时间这一关键维度。现实中,答案正确性常依赖于时间背景。为此,我们提出一个新框架与数据集,涵盖2018至2024年间超过8,000个事件,具有日级粒度,并覆盖政治、科学、商业等全球领域。通过TimeShift评估方法系统测试模型的时序推理能力,发现基础模型在时间敏感召回任务中常优于指令微调和合成训练模型。此外,即使大规模模型在处理改写事实时也表现出脆弱性,凸显时序一致性方面的未解挑战。本研究为发展能适应真实世界动态知识的时间感知语言模型提供了重要进展。
原文摘要 · Abstract (English)
Who is the US President? The answer changes depending on when the question is asked. While large language models (LLMs) are evaluated on various reasoning tasks, they often miss a crucial dimension: time. In real-world scenarios, the correctness of answers is frequently tied to temporal context. To address this gap, we present a novel framework and dataset spanning over 8,000 events from 2018 to 2024, annotated with day-level granularity and sourced globally across domains such as politics, science, and business. Our TimeShift evaluation method systematically probes LLMs for temporal reasoning, revealing that base models often outperform instruction-tuned and synthetic-trained counterparts on time-sensitive recall. Additionally, we find that even large-scale models exhibit brittleness in handling paraphrased facts, highlighting unresolved challenges in temporal consistency. By identifying these limitations, our work provides a significant step toward advancing time-aware language models capable of adapting to the dynamic nature of real-world knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。