arXiv:2506.04410cs.AIcond-mat.mtrl-sci2025-06EMNLP被引 6

构建材料科学假说可行性验证基准,助力自动化科研高效筛选可行假设。

Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science

  • 提出基于回溯验证的可行性评估框架,将假说验证转化为时间过滤任务。
  • 包含8400条跨四大领域的材料科学主张,涵盖理论、实验与仿真结果。
  • 当前模型最佳表现仅72%,远低于专家可解率,凸显提升空间与应用价值。

当前辅助科学发现的方法利用语言模型自动生成大量待测试假说,并自动编写代码实验进行验证。虽然假说生成成本较低,但自动化实验在大规模运行时(如数千次)代价高昂。若能提前评估假说可行性,将显著提升发现系统的效率与成功率。本文提出Matter-of-Fact,一个用于判断科学文献支持假说可行性的挑战性数据集,将可行性评估建模为基于时间过滤的回溯验证任务。该数据集包含从四类高影响力材料科学主题(超导体、半导体、电池、航空航天材料)的论文中提取的8400条主张,涵盖理论、实验及代码/仿真结果中的定性与定量陈述。我们发现,采用文献检索增强生成与代码生成的强基线模型在该任务上最高仅达72%准确率(随机水平为50%),而领域专家认为几乎所有主张均可解决,表明当前模型仍有巨大提升空间,也揭示通过短期进展加速科学发现的巨大潜力。

原文摘要 · Abstract (English)

Contemporary approaches to assisted scientific discovery use language models to automatically generate large numbers of potential hypothesis to test, while also automatically generating code-based experiments to test those hypotheses. While hypotheses can be comparatively inexpensive to generate, automated experiments can be costly, particularly when run at scale (i.e. thousands of experiments). Developing the capacity to filter hypotheses based on their feasibility would allow discovery systems to run at scale, while increasing their likelihood of making significant discoveries. In this work we introduce Matter-of-Fact, a challenge dataset for determining the feasibility of hypotheses framed as claims, while operationalizing feasibility assessment as a temporally-filtered claim verification task using backtesting. Matter-of-Fact includes 8.4k claims extracted from scientific articles spanning four high-impact contemporary materials science topics, including superconductors, semiconductors, batteries, and aerospace materials, while including qualitative and quantitative claims from theoretical, experimental, and code/simulation results. We show that strong baselines that include retrieval augmented generation over scientific literature and code generation fail to exceed 72% performance on this task (chance performance is 50%), while domain-expert verification suggests nearly all are solvable -- highlighting both the difficulty of this task for current models, and the potential to accelerate scientific discovery by making near-term progress.

材料科学假说验证AI科研基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。