将材料的多种数据模态对齐到统一空间,实现跨模态零样本检索。
MatBind: A Shared Embedding Space for Multimodal Materials Characterization

- 以晶体结构为锚点,用对比学习对齐四种材料模态。
- 无需显式配对训练数据,即可实现跨模态零样本检索。
- 适合材料科学中的多模态数据融合与智能检索任务。
完整表征晶态材料需整合异构数据源——原子结构、衍射图谱、电子态密度以及自然语言描述,每种模态捕捉同一物理对象的不同方面。然而,实践中这些模态常被孤立存储与分析,难以跨表示边界关联或查询。本文提出MatBind,一种基于对比学习的框架,以晶体结构为物理锚点,将晶体结构、从结构模拟得到的粉末X射线衍射(pXRD)、态密度(DOS)和文本四种模态对齐至统一嵌入空间。该框架在训练中未显式配对的模态间也实现对齐,从而自然产生零样本跨模态检索能力。所学嵌入空间能按物理意义组织材料,且在查询时融合多个模态可系统性提升检索性能。结果表明,将异构材料数据视为单一物理现实的互补投影,不仅可行,更符合底层物理规律。
原文摘要 · Abstract (English)
Fully characterizing a crystalline material requires integrating heterogeneous data sources -- atomic structures, diffraction patterns, electronic density of states, and natural language -- each of which captures a different facet of the same physical object. In practice, however, these modalities are stored and analyzed in isolation, making it difficult to relate or query materials across representational boundaries. We present MatBind, a contrastive learning framework that aligns four materials modalities -- crystal structure, powder X-ray diffraction (pXRD) simulated from structures, density of states (DOS), and text -- into a unified embedding space using crystal structure as the central physical anchor. The framework induces alignment between modalities never explicitly paired during training, enabling emergent zero-shot cross-modal retrieval as a direct consequence of the shared representation. The learned embedding space organizes materials according to physically meaningful properties without explicit supervision, and retrieval performance improves systematically when modalities are combined at query time. These results demonstrate that treating heterogeneous materials data as complementary projections of a single physical reality, rather than as isolated data sources, is not a practical choice but is consistent with the underlying physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。