arXiv:2503.04150cs.CLcs.AI2025-03被引 1

用干支纪年+极坐标编码,解决大模型时间对齐难题

Temporal Alignment of LLMs through Cycle Encoding for Long-Range Time Representations

  • 改用60年干支周期替代格里高利历,实现时间分布均匀化
  • 极坐标建模干支与年序关系,提升长时序表征能力
  • 适用于需要跨千年时间推理的任务,如历史事件分析

大语言模型在长时序场景下存在时间对齐问题,因其训练数据中时间信息稀疏,导致长期记忆不足或灾难性遗忘。本文提出名为'Ticktack'的方法,在年份粒度上解决该问题。首先,采用60年一循环的干支纪年代替格里高利历,实现更均匀的时间分布;其次,利用极坐标建模干支周期内60个年份及其顺序关系,并引入额外时间编码以增强模型理解;最后,设计一种后训练阶段的时间表征对齐方法,有效区分具有相关知识的时间点,显著提升长时序任务表现。同时构建了覆盖数千年的长时序基准测试集。实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) suffer from temporal misalignment issues especially across long span of time. The issue arises from knowing that LLMs are trained on large amounts of data where temporal information is rather sparse over long times, such as thousands of years, resulting in insufficient learning or catastrophic forgetting by the LLMs. This paper proposes a methodology named "Ticktack" for addressing the LLM's long-time span misalignment in a yearly setting. Specifically, we first propose to utilize the sexagenary year expression instead of the Gregorian year expression employed by LLMs, achieving a more uniform distribution in yearly granularity. Then, we employ polar coordinates to model the sexagenary cycle of 60 terms and the year order within each term, with additional temporal encoding to ensure LLMs understand them. Finally, we present a temporal representational alignment approach for post-training LLMs that effectively distinguishes time points with relevant knowledge, hence improving performance on time-related tasks, particularly over a long period. We also create a long time span benchmark for evaluation. Experimental results prove the effectiveness of our proposal.

时间建模干支纪年长序列大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。