从零散文本复用中重构18世纪论文级重印关系,提升历史文献溯源精度。
Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers
- 通过规则流程整合碎片化复用证据,构建可信的文本传播关系
- 在完整数据集上,最终流程仅标记771对重印,远低于直接大模型的14,886对
- 人工验证确认176个预测为真实重印,适合历史文献学者精准筛选
本文研究从零散文本复用证据中恢复18世纪论文级重印与再利用关系,核心挑战在于成对证据的整合而非片段检索。以苏格兰哲学家大卫·休谟的散文为中心,涵盖ECCO(十八世纪典籍在线)书籍与历史报纸。由于输入为碎片化复用结果且正例覆盖不全,任务被定义为成对证据整合为合理传播关系。对比三种方法:分阶段规则工作流、基线(决策树及两种直接LLM设置)、自动化规则适配。在标注的ECCO-ECCO子集上,仅使用成对特征聚合即可达0.948 F1;最终工作流在整体精度-召回权衡上表现最优。在全量ECCO-ECCO候选集上,直接LLM基线识别出最多14,886对重印,而最终流程仅771对,显示其作为高召回候选扩展器的潜力。在ECCO-报纸数据上,人工审计确认全部176个预测正例均为真实重印或再利用,同时揭示了版本重复与源端多重性带来的额外传记结构。在真值不完整的情况下,可审计的成对证据整合为历史检视提供了紧凑有效的候选空间。
原文摘要 · Abstract (English)
This paper addresses the recovery of essay-scale republication and reuse from fragmented text-reuse evidence, a setting whose central challenge is pair-level evidence consolidation and not fragment retrieval alone. The study focuses on a candidate set centered on essays by eighteenth-century Scottish philosopher David Hume, spanning books from ECCO (Eighteenth Century Collections Online) and historical newspapers. Because the input consists of fragmented reuse hits instead of clean document pairs, and positive coverage is inherently incomplete, we formulate the task as pair-level evidence consolidation into plausible transmission relations and compare three methodological families: a staged rule-based workflow, baselines (a decision tree and two direct LLM settings), and automated rule adaptation. On labeled ECCO--ECCO slices, pair-level feature aggregation alone already reaches 0.948 F1 on the main labeled slice, while the final workflow gives the strongest overall precision-recall trade-off among the tested rule stages. On the full ECCO--ECCO candidate universe, direct LLM baselines flag up to 14,886 pairs as reprints compared to 771 for the final workflow, behaving in this direct-prompt setup as high-recall candidate expanders rather than precision-controlled deployment classifiers. On ECCO--Newspaper, manual audit confirms all 176 predicted positives as genuine cases of republication or reuse, while issue duplication and source-side multiplicity reveal additional provenance structure. Under incomplete ground truth, auditable pair-level evidence consolidation provides a practical way to produce compact candidate spaces for historical inspection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。