arXiv:2506.03822cs.CLcs.IR2025-06中稿 · SCOLIA 2025

构建跨格式网页文档的鲁棒排序数据集,提升文献元数据提取能力

CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents

  • 基于统一嵌入表征融合页面、链接和锚文本信息
  • 在600篇计算机科学论文上实现跨出版商布局无关的准确排序
  • 适合需要从异构网页中提取可靠文献信息的研究者

学术数据库依赖从多样网络来源准确提取元数据,但网页布局与数据格式的差异给元数据提供方带来挑战。本文提出CRAWLDoc,一种用于关联网页文档的上下文排序方法。以出版物的URL(如数字对象标识符)为起点,CRAWLDoc获取着陆页及所有相关资源,包括PDF、ORCID个人资料和补充材料。它将这些资源、锚文本和URL统一嵌入表示。为评估该方法,我们构建了一个由六家顶级计算机科学出版社提供的600篇论文组成的全新人工标注数据集。CRAWLDoc在不同出版商和数据格式下均表现出鲁棒且布局无关的文档相关性排序能力,为从多种布局和格式的网页文档中提升元数据提取奠定了基础。代码与数据集可在https://github.com/FKarl/CRAWLDoc 获取。

原文摘要 · Abstract (English)

Publication databases rely on accurate metadata extraction from diverse web sources, yet variations in web layouts and data formats present challenges for metadata providers. This paper introduces CRAWLDoc, a new method for contextual ranking of linked web documents. Starting with a publication's URL, such as a digital object identifier, CRAWLDoc retrieves the landing page and all linked web resources, including PDFs, ORCID profiles, and supplementary materials. It embeds these resources, along with anchor texts and the URLs, into a unified representation. For evaluating CRAWLDoc, we have created a new, manually labeled dataset of 600 publications from six top publishers in computer science. Our method CRAWLDoc demonstrates a robust and layout-independent ranking of relevant documents across publishers and data formats. It lays the foundation for improved metadata extraction from web documents with various layouts and formats. Our source code and dataset can be accessed at https://github.com/FKarl/CRAWLDoc.

元数据提取文档排序数据集文献挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。