用文本引导联合扩散模型生成周期性材料,性能显著超越现有方法。
Periodic Materials Generation using Text-Guided Joint Diffusion Model
- 通过文本提示在去噪过程中融合全局结构知识,联合生成原子坐标、类型和晶格结构。
- 仅用一个生成样本就超越所有基线模型,结构预测准确率大幅提升。
- 适合材料设计、生成化学家或科研人员使用,尤其擅长基于真实描述生成新材料。
等变扩散模型因其能利用周期性材料结构的物理对称性,已成为生成新型晶体材料的主流方法。然而,现有模型未能在统一端到端框架中有效学习原子类型、分数坐标与晶格结构的联合分布。同时,这些模型均不适用于真实场景——即用户需指定生成结构应满足的特定属性。本文提出TGDMat,一种用于3D周期性材料生成的新型文本引导扩散模型。该方法在每个去噪步骤中引入文本描述以整合全局结构先验,并通过周期性E(3)等变图神经网络联合生成原子坐标、类型及晶格结构。在多个基准数据集上的实验表明,TGDMat显著优于现有基线方法。特别地,在结构预测任务中,仅需一个生成样本即超越所有基线模型,凸显文本引导的重要性;在生成任务中,其性能亦全面领先于所有基线及其文本融合变体,验证了联合扩散范式的有效性。此外,融入文本信息可降低训练与采样计算开销,同时提升在专家提供的真实文本提示下的生成表现。
原文摘要 · Abstract (English)
Equivariant diffusion models have emerged as the prevailing approach for generating novel crystal materials due to their ability to leverage the physical symmetries of periodic material structures. However, current models do not effectively learn the joint distribution of atom types, fractional coordinates, and lattice structure of the crystal material in a cohesive end-to-end diffusion framework. Also, none of these models work under realistic setups, where users specify the desired characteristics that the generated structures must match. In this work, we introduce TGDMat, a novel text-guided diffusion model designed for 3D periodic material generation. Our approach integrates global structural knowledge through textual descriptions at each denoising step while jointly generating atom coordinates, types, and lattice structure using a periodic-E(3)-equivariant graph neural network (GNN). Extensive experiments using popular datasets on benchmark tasks reveal that TGDMat outperforms existing baseline methods by a good margin. Notably, for the structure prediction task, with just one generated sample, TGDMat outperforms all baseline models, highlighting the importance of text-guided diffusion. Further, in the generation task, TGDMat surpasses all baselines and their text-fusion variants, showcasing the effectiveness of the joint diffusion paradigm. Additionally, incorporating textual knowledge reduces overall training and sampling computational overhead while enhancing generative performance when utilizing real-world textual prompts from experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。