arXiv:2503.08600cs.CL2025-03ACL被引 2

从美国国家科学基金会资助项目中挖掘出280万条科学主张,构建大规模数据集。

NSF-SciFy: Mining the NSF Awards Database for Scientific Claims

  • 通过零样本提示技术,从40万份摘要中自动提取科学主张和研究计划。
  • 在主张与研究计划抽取任务上,微调模型相对提升超100%。
  • 适合从事科学发现追踪、科研验证与元科学研究的学者使用。

我们提出NSF-SciFy,一个从美国国家科学基金会(NSF)资助项目摘要中提取的科学主张与研究计划的综合性数据集。此前的科学主张验证数据集规模和范围有限,而NSF-SciFy包含来自40万份摘要的280万条科学主张,覆盖所有科学与数学领域。我们构建了两个聚焦子集:NSF-SciFy-MatSci(11.4万条材料科学主张)和NSF-SciFy-20K(13.5万条跨五个NSF主任办公室的主张)。采用零样本提示方法,实现科学主张与研究计划的联合提取。我们通过三项下游任务验证其价值:生成非技术性摘要、主张提取、研究计划提取。在数据集上微调语言模型后,在主张和研究计划提取任务中相对提升常超过100%。错误分析显示,提取主张具有高精度但召回率较低,提示方法仍有优化空间。NSF-SciFy为大规模科学主张验证、科学发现追踪与元科学研究开辟新路径。代码与数据见https://github.com/darpa-scify/NSFSciFy。

原文摘要 · Abstract (English)

We introduce NSF-SciFy, a comprehensive dataset of scientific claims and investigation proposals extracted from National Science Foundation award abstracts. While previous scientific claim verification datasets have been limited in size and scope, NSF-SciFy represents a significant advance with 2.8 million claims from 400,000 abstracts spanning all science and mathematics disciplines. We present two focused subsets: NSF-SciFy-MatSci with 114,000 claims from materials science awards, and NSF-SciFy-20K with 135,000 claims across five NSF directorates. Using zero-shot prompting, we develop a scalable approach for joint extraction of scientific claims and investigation proposals. We demonstrate the dataset's utility through three downstream tasks: non-technical abstract generation, claim extraction, and investigation proposal extraction. Fine-tuning language models on our dataset yields substantial improvements, with relative gains often exceeding 100%, particularly for claim and proposal extraction tasks. Our error analysis reveals that extracted claims exhibit high precision but lower recall, suggesting opportunities for further methodological refinement. NSF-SciFy enables new research directions in large-scale claim verification, scientific discovery tracking, and meta-scientific analysis. Code and data are available at https://github.com/darpa-scify/NSFSciFy.

科学发现数据挖掘自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。