针对古籍时间检索难题,构建了基于《春秋》的时序精准检索基准与模型。
ChunQiuTR: Time-Keyed Temporal Retrieval in Classical Chinese Annals

- 提出时间锚定检索框架,结合历法上下文与相对时间偏置建模。
- 在月级年号键上实现比基线高12.3%的时序一致性准确率。
- 适合历史文献数字化、时间敏感型知识问答系统研究者使用。
检索影响语言模型在检索增强生成(RAG)中获取与定位知识的方式。在历史研究中,目标往往不是任意相关段落,而是特定统治月份的精确记录,时序一致性与主题相关性同样重要。这对古典汉语编年史尤其具挑战性,因时间以简略、隐含的非公历年号表达,需从上下文推断,导致语义合理但时序错误的证据频现。本文构建了基于《春秋》及其注释传统的时序锚定检索基准ChunQiuTR,按月级年号组织记录,并引入时序相近干扰项模拟真实检索失败场景。进一步提出CTD(历法时间双编码器),融合基于傅里叶的绝对历法上下文与相对偏移偏差机制。实验显示,在时序键评估下,该模型持续优于强基线语义双编码器,支持时序一致性是历史RAG忠实性的关键前提。代码与数据集已开源。
原文摘要 · Abstract (English)
Retrieval shapes how language models access and ground knowledge in retrieval-augmented generation (RAG). In historical research, the target is often not an arbitrary relevant passage, but the exact record for a specific regnal month, where temporal consistency matters as much as topical relevance. This is especially challenging for Classical Chinese annals, where time is expressed through terse, implicit, non-Gregorian reign phrases that must be interpreted from surrounding context, so semantically plausible evidence can still be temporally invalid. We introduce \textbf{ChunQiuTR}, a time-keyed retrieval benchmark built from the \textit{Spring and Autumn Annals} and its exegetical tradition. ChunQiuTR organizes records by month-level reign keys and includes chrono-near confounders that mirror realistic retrieval failures. We further propose \textbf{CTD} (Calendrical Temporal Dual-encoder), a time-aware dual-encoder that combines Fourier-based absolute calendrical context with relative offset biasing. Experiments show consistent gains over strong semantic dual-encoder baselines under time-keyed evaluation, supporting retrieval-time temporal consistency as a key prerequisite for faithful downstream historical RAG. Our code and datasets are available at \href{https://github.com/xbdxwyh/ChunQiuTR}{\texttt{github.com/xbdxwyh/ChunQiuTR}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。