arXiv:2605.04278cs.CL2026-05

用AI自动从论文中提取材料数据,构建可扩展的科学数据库

Material Database Agent: A Multimodal Agentic Framework for Scientific Literature Mining

论文配图:Material Database Agent: A Multimodal Agentic Framework for Scientific Literature Mining
图 1 · 摘自论文原文
  • 多智能体并行处理论文文本与图表,分步构建子数据库
  • 将单篇论文的实验数据转化为结构化表格,支持大规模集成
  • 适合材料科学领域研究者快速获取文献中的实测数据

材料科学研究依赖于大量科学文献中的结构化与非结构化数据。然而,多数实验细节仍深藏于文本、表格、图表和图像之中。传统数据库构建方式耗时费力且难以扩展。多模态大模型使高效准确地从文本和科学图像中提取信息成为可能。本文提出材料数据库智能体(Material Database Agent, MDA),一个模块化的多智能体系统架构,可将科研论文转化为结构化数据库。MDA接收论文PDF作为输入,将其并行解析为Markdown文件和图像。多个子智能体并行读取这些内容,为每篇论文生成子数据库,最终由一个聚合智能体整合为统一表格数据库。与基于规则或单次流程的方法不同,MDA是专为材料科学文献向数据库转化设计的专用架构。本研究为利用多模态智能体进行科学文献信息提取提供了可行路径,推动下一代基于原始文献的科学数据库建设。

原文摘要 · Abstract (English)

Materials science workflows rely on structured and unstructured data from the vast body of available scientific literature. However, most of the experimental details remain buried in text, tables, graphs and figures. Thus, constructing databases that incorporate this data is a manual, time-consuming, and hard-to-scale process. Multimodal large language models have made it feasible to extract information from text and scientific figures with high speed and accuracy. This opens the possibility of an AI system that can create production-scale material databases. Material Database Agent (MDA) is a modular, multi-agent system architecture for converting research literature into structured databases. MDA accepts article PDFs as input, which are subsequently processed in parallel into markdown files and figures. Multiple sub-agents read these markdown files and figures in parallel to assemble sub-databases for each paper. These sub-databases are then compiled into a single tabular database by an agent. As opposed to using either a rule-based approach or a single-pass pipeline for extracting information, MDA is a specialized architecture for transforming the literature into a database in the field of materials science. More generally, this study provides a basis for positioning multimodal agentic information extraction as a viable means for constructing next-generation scientific databases from the primary literature.

材料科学多模态智能体知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。