arXiv:2511.09411cs.CL2025-11中稿 · AAAI

构建细粒度学术实体关系数据集,助力机器学习研究可复现性分析。

GSAP-ERE: Fine-Grained Scholarly Entity and Relation Extraction Focused on Machine Learning

  • 构建包含63K实体、35K关系的细粒度标注数据集
  • 微调模型在实体识别与关系抽取上分别达80.6%和54.0%准确率
  • 揭示大模型提示策略性能远低于有监督模型,凸显数据集重要性

机器学习与人工智能研究发展迅速。从科学文献中进行信息抽取(IE)可大规模识别研究概念与资源信息,有助于提升对机器学习研究的理解与可复现性。为提取机器学习研究中的细粒度信息(如方法训练与数据使用),我们提出GSAP-ERE。这是一个人工标注的细粒度数据集,包含10类实体和18类语义关系,涵盖100篇机器学习论文全文中的63,000个实体提及与35,000条关系。该数据集支持微调模型自动完成知识图谱构建、大规模计算可复现性监控等下游任务。此外,我们以该数据集作为测试基准,评估大语言模型(LLM)的提示策略效果。结果显示,当前最优的LLM提示方法性能显著低于我们最佳微调基线模型(实体识别:80.6% vs. 44.4%;关系抽取:54.0% vs. 10.1%)。这一差距表明,类似GSAP-ERE的数据集对于推动学术信息抽取研究至关重要。

原文摘要 · Abstract (English)

Research in Machine Learning (ML) and AI evolves rapidly. Information Extraction (IE) from scientific publications enables to identify information about research concepts and resources on a large scale and therefore is a pathway to improve understanding and reproducibility of ML-related research. To extract and connect fine-grained information in ML-related research, e.g. method training and data usage, we introduce GSAP-ERE. It is a manually curated fine-grained dataset with 10 entity types and 18 semantically categorized relation types, containing mentions of 63K entities and 35K relations from the full text of 100 ML publications. We show that our dataset enables fine-tuned models to automatically extract information relevant for downstream tasks ranging from knowledge graph (KG) construction, to monitoring the computational reproducibility of AI research at scale. Additionally, we use our dataset as a test suite to explore prompting strategies for IE using Large Language Models (LLM). We observe that the performance of state-of-the-art LLM prompting methods is largely outperformed by our best fine-tuned baseline model (NER: 80.6%, RE: 54.0% for the fine-tuned model vs. NER: 44.4%, RE: 10.1% for the LLM). This disparity of performance between supervised models and unsupervised usage of LLMs suggests datasets like GSAP-ERE are needed to advance research in the domain of scholarly information extraction.

信息抽取知识图谱大模型评测可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。