用自举方法从零构建跨文档细粒度链接数据集,提升链接准确率。
ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
- 通过生成半合成数据集验证并筛选最佳链接方法。
- 结合检索模型与大模型,人工审批通过率达73%。
- 适用于新闻、同行评审等场景,支持后续分析任务。
理解文档间的细粒度关联对诸多应用至关重要,但受限于高效的数据整理方法。为此,我们提出一种无领域依赖的自举框架,可从零开始构建句子级跨文档链接数据集。该方法首先生成并验证半合成的链接文档数据集,其次利用这些数据集评估并筛选出表现最佳的链接方法,最后在大规模人机协同标注中应用优选方法处理真实文本对。我们在同行评审和新闻两个不同领域中应用该框架,发现将检索模型与大语言模型结合,所提链接的人工审批通过率达到73%,超过仅使用强检索器的两倍以上。该框架使用户能够创建新数据集,系统研究跨文档理解,支持媒体框架分析、同行评审评估等下游任务。所有代码、数据及标注协议均已开源,以推动后续研究。
原文摘要 · Abstract (English)
Understanding fine-grained links between documents is crucial for many applications, yet progress is limited by the lack of efficient methods for data curation. To address this limitation, we introduce a domain-agnostic framework for bootstrapping sentence-level cross-document links from scratch. Our approach (1) generates and validates semi-synthetic datasets of linked documents, (2) uses these datasets to benchmark and shortlist the best-performing linking approaches, and (3) applies the shortlisted methods in large-scale human-in-the-loop annotation of natural text pairs. We apply the framework in two distinct domains -- peer review and news -- and show that combining retrieval models with LLMs achieves a 73% human approval rate for suggested links, more than doubling the acceptance of strong retrievers alone. Our framework allows users to produce novel datasets that enable systematic study of cross-document understanding, supporting downstream tasks such as media framing analysis and peer review assessment. All code, data, and annotation protocols are released to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。