首个面向金属有机框架材料的多模态大模型,融合结构与语言理解。
L^2M^3OF: A Large Language Multimodal Model for Metal-Organic Frameworks
- 用晶体编码器将三维结构转为语言可处理的向量表示。
- 在性质预测和知识生成上超越参数更多闭源大模型。
- 适合材料设计、人工智能辅助科研人员使用。
大语言模型在自然语言任务中展现出卓越推理能力,但在科学发现领域进展有限,因复杂物理现象需超越语言的多维度表征。以金属有机框架(MOFs)为例,其在碳捕获、氢能储存等应用中至关重要,但其庞大的三维原子排列空间与严格的配位几何和拓扑规则,使仅用语言描述难以有效建模。尽管早期研究在简单材料系统中取得进展,MOF设计仍高度依赖难以文本化的隐性经验。为此,我们提出L2M3OF,首个面向MOFs的多模态大模型。它结合晶体表征学习与语言理解,联合处理结构、文本与知识模态。采用预训练晶体编码器配合轻量投影层,将结构信息压缩至令牌空间,实现与语言指令的高效对齐。为支持训练与评估,我们构建了晶体材料的结构-性质-知识数据库,并在GPT-5、Gemini-2.5-Pro、DeepSeek-R1等主流闭源大模型上进行基准测试。实验表明,尽管参数量远低于这些模型,L2M3OF在性质预测与知识生成任务中表现更优。结果凸显多模态方法在多孔材料理解中的关键作用,确立了L2M3OF作为下一代材料发现AI系统的基础。
原文摘要 · Abstract (English)
Large language models have demonstrated remarkable reasoning capabilities across diverse natural language tasks. However, comparable breakthroughs in scientific discovery are more limited, because understanding complex physical phenomena demands multifaceted representations far beyond language alone. A compelling example is the design of functional materials such as MOFs-critical for a range of impactful applications like carbon capture and hydrogen storage. Navigating their vast and intricate design space in language-based representations interpretable by LLMs is challenging due to the numerous possible three-dimensional atomic arrangements and strict reticular rules of coordination geometry and topology. Despite promising early results in LLM-assisted discovery for simpler materials systems, MOF design remains heavily reliant on tacit human expertise rarely codified in textual information alone. To overcome this barrier, we introduce L2M3OF, the first multimodal LLM for MOFs. L2M3OF integrates crystal representation learning with language understanding to process structural, textual, and knowledge modalities jointly. L2M3OF employs a pre-trained crystal encoder with a lightweight projection layer to compress structural information into a token space, enabling efficient alignment with language instructions. To facilitate training and evaluation, we curate a structure-property-knowledge database of crystalline materials and benchmark L2M3OF against state-of-the-art closed-source LLMs such as GPT-5, Gemini-2.5-Pro and DeepSeek-R1. Experiments show that L2M3OF outperforms leading text-based closed-source LLMs in property prediction and knowledge generation tasks, despite using far fewer parameters. These results highlight the importance of multimodal approaches for porous material understanding and establish L2M3OF as a foundation for next-generation AI systems in materials discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。