arXiv:2507.01054cs.LGcond-mat.mtrl-sci2025-07

无需晶体结构,仅用元素和XRD数据实现材料性能预测

XxaCT-NN: Structure Agnostic Multimodal Learning for Materials Science

  • 融合元素组成与XRD数据的多模态框架,不依赖晶体结构
  • 在500万样本上预训练,收敛速度提升4.2倍,精度更高
  • 适合实验数据丰富但结构未知的材料研发场景

材料发现的进展主要依赖基于晶体结构的模型,尤其是晶体图模型。然而,这类模型在真实实验中受限于原子结构难以获取。本文提出一种可扩展的多模态框架,直接从元素组成和X射线衍射(XRD)数据中学习,这两者在实验流程中更易获得,无需晶体结构输入。该架构整合了模态专用编码器与交叉注意力融合模块,在500万样本的Alexandria数据集上进行训练。提出掩码XRD建模(MXM)与对比对齐作为自监督预训练策略。预训练使收敛速度提升最多4.2倍,并提高准确率与表征质量。进一步证明,多模态模型性能随数据规模增长更优,大样本下优势持续累积。结果为材料科学中无结构、实验驱动的基础模型提供了可行路径。

原文摘要 · Abstract (English)

Recent advances in materials discovery have been driven by structure-based models, particularly those using crystal graphs. While effective for computational datasets, these models are impractical for real-world applications where atomic structures are often unknown or difficult to obtain. We propose a scalable multimodal framework that learns directly from elemental composition and X-ray diffraction (XRD) -- two of the more available modalities in experimental workflows without requiring crystal structure input. Our architecture integrates modality-specific encoders with a cross-attention fusion module and is trained on the 5-million-sample Alexandria dataset. We present masked XRD modeling (MXM), and apply MXM and contrastive alignment as self-supervised pretraining strategies. Pretraining yields faster convergence (up to 4.2x speedup) and improves both accuracy and representation quality. We further demonstrate that multimodal performance scales more favorably with dataset size than unimodal baselines, with gains compounding at larger data regimes. Our results establish a path toward structure-free, experimentally grounded foundation models for materials science.

材料科学多模态学习结构无关自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。