轻量级模型高效提取学术论文元数据,跨领域泛化能力强。
MeXtract: Light-Weight Metadata Extraction from Scientific Papers
- 基于Qwen 2.5微调0.5B~3B参数模型,专用于论文元数据抽取。
- 在MOLE基准上达当前最优,且对未见模式仍有良好迁移能力。
- 开源代码、数据与模型,支持科研复现与扩展应用。
元数据在科学文献的索引、归档与分析中至关重要,但准确高效地提取仍具挑战。传统方法依赖规则或特定任务模型,难以跨领域和模式变化泛化。本文提出MeXtract,一系列轻量级语言模型,通过微调Qwen 2.5(参数量0.5B至3B)实现科学论文元数据抽取。在相同规模模型中,MeXtract在MOLE基准上达到顶尖性能。为提升评估能力,我们扩展了MOLE基准,加入模型特定元数据,构建更具挑战性的域外子集。实验表明,针对特定模式微调不仅能获得高精度,还能有效迁移到未见过的模式,验证了方法的鲁棒性与适应性。所有代码、数据集与模型均公开发布,供研究社区使用。
原文摘要 · Abstract (English)
Metadata plays a critical role in indexing, documenting, and analyzing scientific literature, yet extracting it accurately and efficiently remains a challenging task. Traditional approaches often rely on rule-based or task-specific models, which struggle to generalize across domains and schema variations. In this paper, we present MeXtract, a family of lightweight language models designed for metadata extraction from scientific papers. The models, ranging from 0.5B to 3B parameters, are built by fine-tuning Qwen 2.5 counterparts. In their size family, MeXtract achieves state-of-the-art performance on metadata extraction on the MOLE benchmark. To further support evaluation, we extend the MOLE benchmark to incorporate model-specific metadata, providing an out-of-domain challenging subset. Our experiments show that fine-tuning on a given schema not only yields high accuracy but also transfers effectively to unseen schemas, demonstrating the robustness and adaptability of our approach. We release all the code, datasets, and models openly for the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。