arXiv:2505.19800cs.CL2025-05EMNLP被引 10

用大模型自动提取科学论文元数据,提升研究可复现性。

MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs

  • 基于提示工程与模式驱动框架,支持多格式文档的自动化元数据提取。
  • 在新基准上验证,大模型在少样本条件下表现良好,上下文长度影响显著。
  • 适合需要高效处理海量论文元数据的研究者或数据平台团队。

元数据提取对数据集的归档、保存以及促进有效科研发现和可复现性至关重要,尤其在科学研究呈指数增长的背景下。尽管Masader(Alyafeai等,2021)为从阿拉伯语NLP数据集的学术论文中提取广泛元数据属性奠定了基础,但其严重依赖人工标注。本文提出MOLE框架,利用大语言模型(LLMs)自动提取涵盖非阿拉伯语语言数据集的科学论文中的元数据属性。该方案采用模式驱动的方法,处理多种输入格式的完整文档,并引入稳健的验证机制以确保输出一致性。此外,我们构建了一个新基准,用于评估该任务的研究进展。通过系统分析上下文长度、少样本学习及网络浏览集成的影响,我们证明现代大模型在自动化该任务方面展现出良好前景,凸显了未来需进一步优化以实现稳定可靠性能的必要性。代码已开源:https://github.com/IVUL-KAUST/MOLE,数据集发布于:https://huggingface.co/datasets/IVUL-KAUST/MOLE,供研究社区使用。

原文摘要 · Abstract (English)

Metadata extraction is essential for cataloging and preserving datasets, enabling effective research discovery and reproducibility, especially given the current exponential growth in scientific research. While Masader (Alyafeai et al.,2021) laid the groundwork for extracting a wide range of metadata attributes from Arabic NLP datasets' scholarly articles, it relies heavily on manual annotation. In this paper, we present MOLE, a framework that leverages Large Language Models (LLMs) to automatically extract metadata attributes from scientific papers covering datasets of languages other than Arabic. Our schema-driven methodology processes entire documents across multiple input formats and incorporates robust validation mechanisms for consistent output. Additionally, we introduce a new benchmark to evaluate the research progress on this task. Through systematic analysis of context length, few-shot learning, and web browsing integration, we demonstrate that modern LLMs show promising results in automating this task, highlighting the need for further future work improvements to ensure consistent and reliable performance. We release the code: https://github.com/IVUL-KAUST/MOLE and dataset: https://huggingface.co/datasets/IVUL-KAUST/MOLE for the research community.

元数据提取大模型应用科研自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。