arXiv:2512.23504cs.CL2025-12

提出新算法自动识别塔木德文献中复杂的引文模式。

Automatic Detection of Complex Quotation Patterns in Aggadic Literature

  • 分三阶段处理,结合形态分析与上下文增强识别嵌套、改写引文。
  • F1达0.91,召回率0.89,精度0.94,优于现有系统。
  • 可识别‘波浪’‘回声’等复杂引文,适合文本分析研究者。

本文提出ACT(Allocate Connections between Texts)算法,用于自动检测拉比文学中的《圣经》引文。与现有文本复用框架不同,该方法针对短句、改写或结构嵌入的引文难题,融合形态感知对齐与上下文敏感增强阶段,识别如“波浪”和“回声”等复杂引用模式。在与Dicta、Passim、Text-Matcher及人工标注权威版本对比中,完整流程ACT-QE表现最优,F1得分为0.91,召回率0.89,精度0.94。三个配置分析显示:缺乏风格增强的ACT-2召回率0.90但精度下降;使用更长n-gram的ACT-3在覆盖与特异性间取得平衡。除提升引文检测外,该模型还具备跨语料库分类风格模式的能力,为体裁分类与互文分析开辟新路径。本工作弥补了机器检测与人工校勘间的差距,为形态丰富、引文密集的传统文本(如阿加迪克文献)提供数字人文与计算考据基础。

原文摘要 · Abstract (English)

This paper presents ACT (Allocate Connections between Texts), a novel three-stage algorithm for the automatic detection of biblical quotations in Rabbinic literature. Unlike existing text reuse frameworks that struggle with short, paraphrased, or structurally embedded quotations, ACT combines a morphology-aware alignment algorithm with a context-sensitive enrichment stage that identifies complex citation patterns such as "Wave" and "Echo" quotations. Our approach was evaluated against leading systems, including Dicta, Passim, Text-Matcher, as well as human-annotated critical editions. We further assessed three ACT configurations to isolate the contribution of each component. Results demonstrate that the full ACT pipeline (ACT-QE) outperforms all baselines, achieving an F1 score of 0.91, with superior Recall (0.89) and Precision (0.94). Notably, ACT-2, which lacks stylistic enrichment, achieves higher Recall (0.90) but suffers in Precision, while ACT-3, using longer n-grams, offers a tradeoff between coverage and specificity. In addition to improving quotation detection, ACT's ability to classify stylistic patterns across corpora opens new avenues for genre classification and intertextual analysis. This work contributes to digital humanities and computational philology by addressing the methodological gap between exhaustive machine-based detection and human editorial judgment. ACT lays a foundation for broader applications in historical textual analysis, especially in morphologically rich and citation-dense traditions like Aggadic literature.

文本检测数字人文引文分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。