用多模态学习统一材料结构、谱图和成分数据,提升嵌入表征效果。
UniMat: Unifying Materials Embeddings through Multi-modal Learning
- 通过对齐与融合原子结构、XRD谱图和成分数据,构建统一嵌入空间。
- 结构图与XRD对齐后性能提升,融合实验易获取数据更鲁棒。
- 适合材料设计与发现中需跨模态分析的研究者使用。
材料科学数据天然异构,涵盖表征谱图、原子结构、显微图像和文本形式的合成条件。多模态学习(尤其是视觉与语言模型)的发展为整合不同形式的数据提供了新路径。本文评估了多模态学习中的常见技术(对齐与融合)在统一材料科学中关键模态——原子结构、X射线衍射图谱(XRD)和化学组成——上的效果。结果表明,将结构图与XRD谱图对齐可增强其表示能力;同时,对齐并融合实验上更易获取的XRD谱图与成分数据,能生成比单一模态更稳健的联合嵌入,在多种任务中表现更优。这为未来充分挖掘材料科学多模态数据潜力奠定了基础,有助于实现更明智的材料设计与发现决策。
原文摘要 · Abstract (English)
Materials science datasets are inherently heterogeneous and are available in different modalities such as characterization spectra, atomic structures, microscopic images, and text-based synthesis conditions. The advancements in multi-modal learning, particularly in vision and language models, have opened new avenues for integrating data in different forms. In this work, we evaluate common techniques in multi-modal learning (alignment and fusion) in unifying some of the most important modalities in materials science: atomic structure, X-ray diffraction patterns (XRD), and composition. We show that structure graph modality can be enhanced by aligning with XRD patterns. Additionally, we show that aligning and fusing more experimentally accessible data formats, such as XRD patterns and compositions, can create more robust joint embeddings than individual modalities across various tasks. This lays the groundwork for future studies aiming to exploit the full potential of multi-modal data in materials science, facilitating more informed decision-making in materials design and discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。