统一多任务分子生成的3D扩散框架,提升精度与泛化能力
MODA: A Unified 3D Diffusion Framework for Multi-Task Target-Aware Molecular Generation
- 基于贝叶斯掩码调度的单阶段多任务训练,共享几何化学先验
- 在子结构、性质、相互作用和几何上超越6个基线模型
- 支持零样本从头设计与先导优化,无需力场微调
基于扩散模型的三维分子生成器现已接近晶体学精度,但任务间仍呈碎片化。仅使用SMILES输入、两阶段预训练微调流程及一任务一模型的做法限制了立体化学保真度、任务对齐性和零样本迁移能力。我们提出MODA,一个统一片段生长、连接子设计、骨架跃迁和侧链修饰的扩散框架,采用贝叶斯掩码调度。训练时,连续空间片段被掩码后单步去噪,使模型学习跨任务共享的几何与化学先验。多任务训练得到的通用主干在子结构、化学性质、相互作用和几何上优于六个扩散基线与三种训练范式。Model-C减少配体-蛋白冲突与子结构偏离,同时保持Lipinski合规;Model-B保持相似性但新颖性与结合亲和力较低。零样本从头设计与先导优化测试表明,无需力场精调即可稳定获得负Vina评分与高提升率。结果表明,单阶段多任务扩散流程可取代传统两阶段结构导向分子设计流程。
原文摘要 · Abstract (English)
Three-dimensional molecular generators based on diffusion models can now reach near-crystallographic accuracy, yet they remain fragmented across tasks. SMILES-only inputs, two-stage pretrain-finetune pipelines, and one-task-one-model practices hinder stereochemical fidelity, task alignment, and zero-shot transfer. We introduce MODA, a diffusion framework that unifies fragment growing, linker design, scaffold hopping, and side-chain decoration with a Bayesian mask scheduler. During training, a contiguous spatial fragment is masked and then denoised in one pass, enabling the model to learn shared geometric and chemical priors across tasks. Multi-task training yields a universal backbone that surpasses six diffusion baselines and three training paradigms on substructure, chemical property, interaction, and geometry. Model-C reduces ligand-protein clashes and substructure divergences while maintaining Lipinski compliance, whereas Model-B preserves similarity but trails in novelty and binding affinity. Zero-shot de novo design and lead-optimisation tests confirm stable negative Vina scores and high improvement rates without force-field refinement. These results demonstrate that a single-stage multi-task diffusion routine can replace two-stage workflows for structure-based molecular design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。