用语言模型直接理解与生成材料结构,打通自然语言与原子构型的桥梁。
Atomistic Language Models Understand and Generate Materials

- 统一语言模型、原子编码器与扩散模型,通过连续投影实现跨模态融合。
- 在晶体结构预测与从头生成任务上达到当前最优性能。
- 适合材料设计、智能化学与多模态生成方向的研究者使用。
原子结构与自然语言长期被分别建模,语言模型要么作为工具调用原子模型,要么在丢失原子信息的文本编码上微调。我们提出原子语言模型(ALMs),以原生多模态为目标,使单一语言主干能够理解原子结构、根据自然语言生成材料,并按文本指令优化晶格结构。通过纯连续投影器与分阶段训练,将预训练原子编码器、大语言模型和去噪扩散模型统一起来,ALMs在晶体结构预测与从头生成任务上取得当前最优结果。其核心是语言模型嵌入直接映射到原子扩散的控制空间的连续桥梁,并由基于粒子的采样器Text-to-Crystal Feynman-Kac(T2C-FK)辅助,在推理时通过评分部分去噪轨迹来强制满足化学计量目标。为评估ALMs在自然语言提示与3D原子坐标输入下的材料优化与生成能力,我们引入了首个文本条件晶体生成与优化基准——ALM Bench。代码、训练数据与模型权重即将发布。
原文摘要 · Abstract (English)
Atomistic structure and natural language have long been modeled separately, with language models either calling atomistic models as tools or being fine-tuned on lossy textual encodings that discard atomistic information. We introduce Atomistic Language Models (ALMs) to pursue native multimodality, in which a single language backbone understands atomistic structures, generates materials from natural language, and optimizes crystal structures as instructed by text. By unifying a pretrained atomistic encoder, large language model, and denoising diffusion model through purely continuous projectors and staged training, ALMs achieve state-of-the-art results on crystal structure prediction and de novo generation. ALMs are enabled by a continuous bridge that maps language model embeddings directly into the steering space of atomistic diffusion, and are assisted by Text-to-Crystal Feynman-Kac (T2C-FK), a particle-based sampler that scores partial denoising trajectories to enforce stoichiometric targets at inference time. To evaluate the ability of ALMs to optimize and generate materials from natural-language prompts and 3D atom-coordinate inputs, we introduce ALM Bench, the first benchmark for text-conditioned crystal generation and optimization. Code, training data, and model weights will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。