arXiv:2508.13251cs.AIcond-mat.mtrl-sci2025-08被引 22

用AI从论文图表中自动提取氢储材料数据,加速新材料发现。

"DIVE" into Hydrogen Storage Materials Discovery with AI Agents

  • 设计多智能体流程DIVE,从图文混排的论文中解析实验数据
  • 相比开源模型提升超30%数据提取准确率,覆盖率达3万+条记录
  • 2分钟内完成全新储氢材料逆向设计,适用于各类功能材料

数据驱动的人工智能正重塑新材料发现范式。尽管科学文献中材料数据海量,但多数信息仍困于非结构化的图表中,限制了基于大语言模型(LLM)的AI智能体在自动化材料设计中的应用。本文提出描述性视觉表达解析(DIVE)多智能体工作流,系统化读取并组织科学文献中的图形元素数据。聚焦固态储氢材料——未来清洁能源技术的核心材料类别,DIVE显著提升了数据提取的准确性和覆盖率:相较于直接使用多模态模型,商业模型提升10-15%,开源模型提升超30%。基于从4,000篇文献中整理的超3万条数据,构建快速逆向设计流程,可在两分钟内识别出此前未报道的储氢材料组成。所提出的AI工作流与智能体设计具有广泛可迁移性,为人工智能驱动的新材料发现提供了新范式。

原文摘要 · Abstract (English)

Data-driven artificial intelligence (AI) approaches are fundamentally transforming the discovery of new materials. Despite the unprecedented availability of materials data in the scientific literature, much of this information remains trapped in unstructured figures and tables, hindering the construction of large language model (LLM)-based AI agent for automated materials design. Here, we present the Descriptive Interpretation of Visual Expression (DIVE) multi-agent workflow, which systematically reads and organizes experimental data from graphical elements in scientific literatures. We focus on solid-state hydrogen storage materials-a class of materials central to future clean-energy technologies and demonstrate that DIVE markedly improves the accuracy and coverage of data extraction compared to the direct extraction by multimodal models, with gains of 10-15% over commercial models and over 30% relative to open-source models. Building on a curated database of over 30,000 entries from 4,000 publications, we establish a rapid inverse design workflow capable of identifying previously unreported hydrogen storage compositions in two minutes. The proposed AI workflow and agent design are broadly transferable across diverse materials, providing a paradigm for AI-driven materials discovery.

AI发现材料氢储能多智能体数据提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。