构建多模态材料数据集,让机器理解晶体结构与文本描述
Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework
- 将原子结构、二维投影和文本标注结合,形成多模态数据集
- 支持属性预测与部分监督下的晶体生成,性能优于传统几何数据
- 引入人机协作流程,提升标注质量,适合材料与AI交叉研究者
多数材料科学数据集仅包含原子几何信息(如XYZ文件),限制了多模态学习与数据驱动分析的应用。本文提出多模态晶体谱集(MCS-Set),通过整合原子结构、二维投影及结构化文本注释(包括晶格参数与配位度量),拓展材料数据集。MCS-Set支持两项关键任务:(1) 多模态属性与摘要预测,(2) 基于部分聚类监督的晶体生成。采用人机协作流程,融合领域知识与标准化描述符,实现高质量标注。基于先进语言与视觉-语言模型的评估显示显著的模态性能差异,凸显标注质量对泛化能力的重要性。该框架为多模态模型基准测试、标注实践改进及可访问、多功能材料数据集提供基础。数据集与代码已开源:https://github.com/KurbanIntelligenceLab/MultiCrystalSpectrumSet。
原文摘要 · Abstract (English)
Most materials science datasets are limited to atomic geometries (e.g., XYZ files), restricting their utility for multimodal learning and comprehensive data-centric analysis. These constraints have historically impeded the adoption of advanced machine learning techniques in the field. This work introduces MultiCrystalSpectrumSet (MCS-Set), a curated framework that expands materials datasets by integrating atomic structures with 2D projections and structured textual annotations, including lattice parameters and coordination metrics. MCS-Set enables two key tasks: (1) multimodal property and summary prediction, and (2) constrained crystal generation with partial cluster supervision. Leveraging a human-in-the-loop pipeline, MCS-Set combines domain expertise with standardized descriptors for high-quality annotation. Evaluations using state-of-the-art language and vision-language models reveal substantial modality-specific performance gaps and highlight the importance of annotation quality for generalization. MCS-Set offers a foundation for benchmarking multimodal models, advancing annotation practices, and promoting accessible, versatile materials science datasets. The dataset and implementations are available at https://github.com/KurbanIntelligenceLab/MultiCrystalSpectrumSet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。