arXiv:2501.05082cs.IRcs.CL2025-01

对比多种方法从模板多样的学术文档中提取元数据

Comparison of Feature Learning Methods for Metadata Extraction from PDF Scholarly Documents

  • 融合NLP、计算机视觉与多模态方法提取元数据
  • 在模板差异大的文档中仍保持较高提取准确率
  • 适合关注科学文献可发现性与可重用性的研究者

科学文献的元数据对于推动科研进展及遵循研究结果的FAIR原则(可发现性、可访问性、互操作性、可重用性)至关重要。然而,许多小型和中型出版商发布的文献缺乏足够的元数据,尤其在德国社会科学领域,因使用多样化的排版模板,这一问题尤为突出。为此,本研究评估了多种特征学习与预测方法,包括自然语言处理(NLP)、计算机视觉(CV)及多模态方法,在高模板变异性文档中提取元数据的性能。实验提供了全面的结果,分析了各方法在准确性与效率方面的表现,并揭示了各类方法的优缺点,为该领域的后续研究提供指导。

原文摘要 · Abstract (English)

The availability of metadata for scientific documents is pivotal in propelling scientific knowledge forward and for adhering to the FAIR principles (i.e. Findability, Accessibility, Interoperability, and Reusability) of research findings. However, the lack of sufficient metadata in published documents, particularly those from smaller and mid-sized publishers, hinders their accessibility. This issue is widespread in some disciplines, such as the German Social Sciences, where publications often employ diverse templates. To address this challenge, our study evaluates various feature learning and prediction methods, including natural language processing (NLP), computer vision (CV), and multimodal approaches, for extracting metadata from documents with high template variance. We aim to improve the accessibility of scientific documents and facilitate their wider use. To support our comparison of these methods, we provide comprehensive experimental results, analyzing their accuracy and efficiency in extracting metadata. Additionally, we provide valuable insights into the strengths and weaknesses of various feature learning and prediction methods, which can guide future research in this field.

元数据提取NLP多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。