arXiv:2607.23886cond-mat.mtrl-scics.AI2026-07

将电池文献中的X射线谱图转化为可分析的数据集,助力材料发现。

Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature

论文配图:Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature
图 1 · 摘自论文原文
  • 通过图文联合分析识别文献中的XAS谱图并数字化曲线。
  • 构建13740条谱图的开放数据集,覆盖66种元素和多种电池体系。
  • 适合材料科学、电池研发及人工智能驱动发现的研究者使用。

X射线吸收谱(XAS)是理解材料局部电子与原子结构的关键手段,但多数已发表谱图因嵌入图片且文本描述零散,难以用于数据驱动分析。本文提出一种多模态(图像与文本)文献挖掘方法,将分散的知识转化为面向AI的实验数据资源。开发了可扩展的谱图数字化流程,能自动识别全文中XAS图像,提取光谱曲线,并关联吸收边与材料元数据。该流程应用于电池领域文献,生成包含13,740条XAS谱图的开放数据集,涵盖66种吸收元素和多种电池化学体系,经专家验证,谱图与元数据提取准确。通过将文献中的谱图转为结构化数值数据,该数据集为大规模XAS分析、跨实验室比较、高通量表征及先进材料的自主发现提供了基础。

原文摘要 · Abstract (English)

X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature. Here, we use multimodal (image and text) literature mining to transform this dispersed knowledge into an AI-ready experimental data resource. We developed a scalable spectroscopy data digitization pipeline that identifies XAS figures in full-text articles, digitizes spectral curves, and links each spectrum to accompanying metadata on the measured edge and material. Applying this pipeline to the battery literature produced an open dataset of 13,740 XAS spectra, spanning 66 absorbing elements and diverse battery chemistries, with expert validation confirming accurate extraction of spectral and metadata information. By converting literature-embedded spectra into structured numerical data, this dataset provides a foundation for large-scale XAS analysis, cross-laboratory comparison, high-throughput characterization, and autonomous discovery of advanced materials.

XAS电池材料数据挖掘多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。