arXiv:2607.27595cs.CLcs.AI2026-07

用大模型精准识别古籍间的引用细节,揭示文本传承的深层规律。

Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories

论文配图:Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
图 1 · 摘自论文原文
  • 让大模型扮演分析员,精确定位引用位置并分类为五维度标签。
  • 在2533对古籍引用中验证,模型准确率56%至93%,成本相差51倍。
  • 发现引用整体稳定但逐字引用减少,反映文化传承的动态平衡。

计算方法从字符串匹配发展到神经检索,但仅能给出相似度分数和对应段落列表,无法解释引用方式与动机。本文将细粒度互文性提取重构为代理任务:大语言模型需完整阅读两段文本,通过受限工具接口精确标注引用字符范围,并按形式、方面、源标记、功能、立场五个维度分类。在《论语》与《汉书》的全面对比中,三位领域专家共同评审多模型候选集,形成包含2533个互文对的基准数据集。基于此标准,评估12个大模型,报告精度56%-93%,相同质量下成本相差51倍,以及置信度校准情况。专家一致性显示:表层可读维度标注一致,需推断意图的维度存在争议,限制了标注结论的适用范围。将验证后的提取器扩展至全《二十四史》(65,380次比较,5,766对互文),揭示了相似度无法表达的语料级结构。引文的解释性构成在十八世纪间无系统变化,但引文的字面忠实度持续下降。整体稳定与个体漂移并存,符合文化吸引理论预期。研究发布提取协议与专家标注基准。

原文摘要 · Abstract (English)

Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how or why. We recast fine-grained intertextuality extraction as an agentic task in which a large language model (LLM) reads two text units in full and, through a constrained tool interface, must ground each proposed reuse in exact character spans on both sides and label it under a five-dimension typology of reuse (form, aspect, source-marking, function, stance). We validate the approach on an exhaustive comparison of the Analects with the Book of Han, where three domain experts adjudicate a pooled multi-model candidate set into a benchmark of 2,533 intertextual pairs. Against this standard we study twelve LLMs, reporting precision (56%-93%), a 51$\times$ cost spread at comparable quality, and how well their confidence is calibrated. Expert agreement traces a reliability gradient: dimensions legible on the textual surface are annotated consistently, while those requiring inference of intent are contested, delimiting the claims such annotation supports. Scaling the validated extractor to the full Twenty-Four Histories (65,380 comparisons, 5,766 pairs) recovers corpus-level structure a similarity score cannot express. The interpretive composition of citation shows no systematic change across eighteen centuries, yet the same passage is quoted less and less literally. Stability in the aggregate with drift in the individual case is what a cultural-attraction account expects. We release the extraction protocol and the expert-adjudicated benchmark.

古籍分析大模型应用互文性文化传承

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。