构建文献实验提取基准,提升材料数据自动化获取精度。
LitXBench: A Benchmark for Extracting Experiments from Scientific Literature

- 以Python对象存储实验数据,增强可审计性和程序验证能力
- 前沿模型在提取准确率上比现有流程高0.37 F1
- 适合材料科学与文献挖掘交叉研究者使用
从论文中聚合实验数据有助于材料科学家构建更优的性能预测模型并推动科学发现。近年来,研究兴趣转向不仅提取单一材料属性,还包括完整实验测量。为此,我们提出LitXBench框架,用于评估从文献中提取实验的方法。同时发布LitXAlloy,一个包含19篇合金论文中总计1426个测量值的密集基准。通过将基准条目以Python对象形式存储,而非传统的CSV或JSON文本格式,提升了可审计性并支持程序化数据验证。结果显示,前沿语言模型(如Gemini 3.1 Pro Preview)在提取任务中表现优于现有多轮提取流水线,最高达0.37 F1提升。分析表明,该性能差距源于现有流水线将测量值关联到成分,而非定义材料的制备步骤。
原文摘要 · Abstract (English)
Aggregating experimental data from papers enables materials scientists to build better property prediction models and to facilitate scientific discovery. Recently, interest has grown in extracting not only single material properties but also entire experimental measurements. To support this shift, we introduce LitXBench, a framework for benchmarking methods that extract experiments from literature. We also present LitXAlloy, a dense benchmark comprising 1426 total measurements from 19 alloy papers. By storing the benchmark's entries as Python objects, rather than text-based formats such as CSV or JSON, we improve auditability and enable programmatic data validation. We find that frontier language models, such as Gemini 3.1 Pro Preview, outperform existing multi-turn extraction pipelines by up to 0.37 F1. Our results suggest that this performance gap arises because extraction pipelines associate measurements with compositions rather than the processing steps that define a material.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。