arXiv:2604.07584cs.AI2026-04被引 3

用大模型自动提取材料实验数据,实现高精度结构化整理。

From Papers to Property Tables: A Priority-Based LLM Workflow for Materials Data Extraction

  • 分三级优先级整合文本、表格、图表和物理公式信息,逐层提取数据。
  • 对30篇论文1.2万条数据测试,整体准确率达94.69%,三级提取准确率分别为94.93%、92.04%、83.49%。
  • 无需微调即可批量处理,适合构建可追溯的材料数据库,尤其适合研究者与工程师使用。

科学数据广泛分散于科研文章中,常以文本、表格和图表形式不一致地呈现,导致手动提取与聚合效率低且易出错。本文提出一种基于提示的层级化工作流,利用大语言模型(LLM)从全文发表的研究文章中自动提取并重构结构化的冲击物理实验记录,以合金溅射强度为典型案例。该流程针对每次实验的37个相关字段,采用三级优先策略:(T1)直接从文本/表格中提取,(T2)基于已验证的物理关系推导,(T3)必要时从图表中数字化获取。提取结果统一归一化至标准单位,按优先级打标以确保可追溯性,并通过物理一致性与合理性检验。在包含30篇论文、共11,967个数据点的基准测试中,整体准确率达到94.69%,各级准确率分别为94.93%(T1)、92.04%(T2)和83.49%(T3)。跨模型测试显示,文本/表格与方程推导字段一致性高,图示提取一致性较低。通过API接口实现,展示了方法的可扩展性,在部分测试中性能达到或超过聊天式交互水平。该工作流为将非结构化技术文献转化为可追踪、可分析的数据集提供了实用路径,无需任务特定微调,支持材料科学领域的大规模数据库建设。

原文摘要 · Abstract (English)

Scientific data are widely dispersed across research articles and are often reported inconsistently across text, tables, and figures, making manual data extraction and aggregation slow and error-prone. We present a prompt-driven, hierarchical workflow that uses a large language model (LLM) to automatically extract and reconstruct structured, shot-level shock-physics experimental records by integrating information distributed across text, tables, figures, and physics-based derivations from full-text published research articles, using alloy spall strength as a representative case study. The pipeline targeted 37 experimentally relevant fields per shot and applied a three-level priority strategy: (T1) direct extraction from text/tables, (T2) physics-based derivation using verified governing relations, and (T3) digitization from figures when necessary. Extracted values were normalized to canonical units, tagged by priority for traceability, and validated with physics-based consistency and plausibility checks. Evaluated on a benchmark of 30 published research articles comprising 11,967 evaluated data points, the workflow achieved high overall accuracy, with priority-wise accuracies of 94.93% (T1), 92.04% (T2), and 83.49% (T3), and an overall weighted accuracy of 94.69%. Cross-model testing further indicated strong agreement for text/table and equation-derived fields, with lower agreement for figure-based extraction. Implementation through an API interface demonstrated the scalability of the approach, achieving consistent extraction performance and, in a subset of test cases, matching or exceeding chat-based accuracy. This workflow demonstrates a practical approach for converting unstructured technical literature into traceable, analysis-ready datasets without task-specific fine-tuning, enabling scalable database construction in materials science.

数据提取大模型应用材料科学自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。