arXiv:2607.29044cs.CL2026-07

将古籍注释按语境自动整理,提升经典文本研究效率

From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts

论文配图:From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts
图 1 · 摘自论文原文
  • 用两步提示链识别注释与正文的关联及解释功能
  • 跨版本聚类实现多版注释整合,CoNLL F1超97%
  • 适合古籍数字化、训诂学与历史文献研究者

inline注释与汇编注疏是儒家训诂传统中重要的学术表达形式,但尚未受到计算关注。本文结合中国传统训诂学与语言学,将汇编注疏定义为自然语言处理任务,提出一种保持注释上下文依赖的计算框架。该框架通过两步提示链识别注释对应的正文段落及其训诂功能,并利用跨源指代聚类整合不同版本的注释,在《山海经》案例研究中取得超过97%的CoNLL F1分数。该方法为大规模历史训诂知识组织奠定基础,支持下游多种训诂与NLP任务。

原文摘要 · Abstract (English)

Inline notes and collected commentaries are important forms of scholarly communication that evolved within the Confucian exegetical tradition, yet have received little computational attention. Drawing on traditional Chinese exegetics and philology, this paper formulates collected commentary compilation as an NLP task and proposes a computational framework that preserves the contextual dependency of inline notes while enabling their automatic compilation and exegetical knowledge organization. It combines two-step prompt chaining for identifying the associated main-text segments and exegetical functions of annotations with cross-source mention clustering for integrating commentary across editions, achieving a CoNLL F1 score above 97% in a case study on the Classic of Mountains. Our framework lays the foundation for the large-scale organization of historical exegetical knowledge, thereby supporting a broad range of downstream philological and NLP tasks.

古籍处理注疏整理NLP应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。