提出MELD模型,让分子扩散更灵活,显著提升生成质量
Learning Flexible Forward Trajectories for Masked Molecular Diffusion
- 为每个原子和键设计独立噪声调度,避免分子间状态冲突
- 在ZINC250K上化学有效性从15%提升至93%,性能大幅跃升
- 适合需要高质量分子生成与属性对齐的研究者使用
掩码扩散模型(MDMs)在离散数据建模中取得显著进展,但在分子生成领域的潜力尚未充分探索。本文发现,直接应用标准MDMs会严重降低性能,根源在于不同分子的前向扩散过程会坍缩到同一状态,导致重建目标混合,无法通过常规单模态反向扩散学习。为此,我们提出掩码元素可学习扩散(MELD),通过参数化噪声调度网络为每个图元素(原子和键)分配独立的破坏率,实现逐元素污染轨迹,避免分子间状态碰撞。在多个分子基准测试上的实验表明,相比无差别噪声调度,MELD显著提升生成质量,在ZINC250K上将原始MDMs的化学有效性从15%提升至93%,并在条件生成任务中达到最先进的属性对齐效果。
原文摘要 · Abstract (English)
Masked diffusion models (MDMs) have achieved notable progress in modeling discrete data, while their potential in molecular generation remains underexplored. In this work, we explore their potential and introduce the surprising result that naively applying standards MDMs severely degrades the performance. We identify the critical cause of this issue as a state-clashing problem-where the forward diffusion of distinct molecules collapse into a common state, resulting in a mixture of reconstruction targets that cannot be learned using typical reverse diffusion process with unimodal predictions. To mitigate this, we propose Masked Element-wise Learnable Diffusion (MELD) that orchestrates per-element corruption trajectories to avoid collision between distinct molecular graphs. This is achieved through a parameterized noise scheduling network that assigns distinct corruption rates to individual graph elements, i.e., atoms and bonds. Extensive experiments on diverse molecular benchmarks reveal that MELD markedly enhances overall generation quality compared to element-agnostic noise scheduling, increasing the chemical validity of vanilla MDMs on ZINC250K from 15% to 93%, Furthermore, it achieves state-of-the-art property alignment in conditional generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。