用可解释机器学习融合光谱与结构数据,解析氧化物中过渡金属的局部环境。
Interpretable Multimodal Machine Learning Analysis of X-ray Absorption Near-Edge Spectra and Pair Distribution Functions
- 用随机森林模型融合XANES和PDF数据,预测氧化态、配位数和键长。
- 仅用XANES就比纯PDF表现更好,尤其在结构任务中,但差分PDF缩小了差距。
- 方法快速易用,适合指导实验设计,判断多模态数据是否真正互补。
我们采用可解释机器学习,融合异构光谱数据——X射线吸收近边谱(XANES)和原子对分布函数(PDF),以提取氧化物中过渡金属阳离子的局部结构与化学环境。通过模拟数据训练随机森林模型,分别针对XANES、PDF及两者联合输入,预测氧化态、配位数和平均最近邻键长。结果显示,仅使用XANES的模型在多数任务上优于仅使用PDF的模型,即使在结构预测中亦然;但改用金属特异性差分PDF(dPDF)后,性能差距显著缩小。当两种数据结合时,XANES信息通常主导预测结果。研究证明XANES包含丰富结构信息,凸显其物种特异性优势。该可解释的多模态方法易于实现,依赖可靠数据库,为科学实验设计提供参考,帮助判断不同表征手段组合是否带来实质性增益。
原文摘要 · Abstract (English)
We used interpretable machine learning to combine information from multiple heterogeneous spectra: X-ray absorption near-edge spectra (XANES) and atomic pair distribution functions (PDFs) to extract local structural and chemical environments of transition metal cations in oxides. Random forest models were trained on simulated XANES, PDF, and both combined to extract oxidation state, coordination number, and mean nearest-neighbor bond length. XANES-only models generally outperformed PDF-only models, even for structural tasks, although using the metal's differential PDFs (dPDFs) instead of total PDFs narrowed this gap. When combined with PDFs, information from XANES often dominates the prediction. Our results demonstrate that XANES contain rich structural information and highlight the utility of species-specificity. This interpretable, multimodal approach is quick to implement with suitable databases and offers valuable insights into the relative strengths of different modalities, guiding researchers in experiment design and identifying when combining complementary techniques adds meaningful information to a scientific investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。