arXiv:2604.04403cs.AI2026-04

用扩散模型生成分子,更准更有效。

MolDA: Molecular Understanding and Generation via Large Language Diffusion Model

  • 用扩散模型替代传统逐字生成,避免结构错误积累。
  • 能生成化学上有效的分子,且保持全局结构一致。
  • 适合药物发现、分子设计等需要高质量分子的场景。

大语言模型在分子发现中取得显著进展,但现有多模态分子架构仍依赖自回归(AR)骨干网络。这种严格的从左到右归纳偏置不利于生成化学有效分子,因难以处理非局部全局约束(如环闭合),且在序列生成过程中容易累积结构错误。为此,我们提出MolDA(带掩码扩散的分子语言模型),用离散大语言扩散模型取代传统AR骨干。MolDA通过混合图编码器提取全面的结构表征,捕捉局部与全局拓扑,并通过Q-Former将它们对齐至语言标记空间。此外,我们为掩码扩散量身定制了分子结构偏好优化的数学重述。通过双向迭代去噪,MolDA确保分子生成、描述和性质预测中的全局结构一致性、化学有效性与强推理能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have significantly advanced molecular discovery, but existing multimodal molecular architectures fundamentally rely on autoregressive (AR) backbones. This strict left-to-right inductive bias is sub-optimal for generating chemically valid molecules, as it struggles to account for non-local global constraints (e.g., ring closures) and often accumulates structural errors during sequential generation. To address these limitations, we propose MolDA (Molecular language model with masked Diffusion with mAsking), a novel multimodal framework that replaces the conventional AR backbone with a discrete Large Language Diffusion Model. MolDA extracts comprehensive structural representations using a hybrid graph encoder, which captures both local and global topologies, and aligns them into the language token space via a Q-Former. Furthermore, we mathematically reformulate Molecular Structure Preference Optimization specifically for the masked diffusion. Through bidirectional iterative denoising, MolDA ensures global structural coherence, chemical validity, and robust reasoning across molecule generation, captioning, and property prediction.

分子生成扩散模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。