一个模型搞定晶体生成各类任务,统一建模更高效。
Multimodal Crystal Flow: Any-to-Any Modality Generation for Unified Crystal Modeling
- 用独立时间变量实现原子类型与晶格结构的联合生成
- 在MP-20和MPTS-52上超越专用模型性能
- 无需模板即可注入化学与对称性先验知识
晶体建模涵盖晶体结构预测(CSP)和从头生成(DNG)等条件与非条件生成任务。尽管深度生成模型表现良好,但大多局限于特定任务,缺乏跨任务共享晶体表示的统一框架。为此,我们提出多模态晶体流(MCFlow),通过为原子类型和晶体结构分别设置独立时间变量,将多种晶体生成任务统一为不同的推理轨迹。为在标准Transformer中实现多模态流,我们引入一种兼顾成分与对称性的原子排序方式,并结合分层排列增强,无需显式结构模板即可注入化学与晶体学先验。在MP-20和MPTS-52基准测试中,单一MCFlow模型在CSP、DNG及结构引导的原子类型生成任务上均达到与专用基线相当甚至更优的表现。
原文摘要 · Abstract (English)
Crystal modeling spans a family of conditional and unconditional generation tasks, including crystal structure prediction (CSP) and de novo generation (DNG). While recent deep generative models have shown promising performance, they remain largely task-specific, lacking a unified framework that shares crystal representations across tasks. To address this limitation, we propose Multimodal Crystal Flow (MCFlow), a unified multimodal flow model that realizes multiple crystal generation tasks as distinct inference trajectories via independent time variables for atom types and crystal structures. To enable multimodal flow in a standard transformer model, we introduce a composition- and symmetry-aware atom ordering with hierarchical permutation augmentation, injecting compositional and crystallographic priors without explicit structural templates. Experiments on the MP-20 and MPTS-52 benchmarks show that a single MCFlow model is competitive with task-specific baselines across CSP, DNG, and structure-conditioned atom type generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。