分步生成分子结构,让复杂描述更准确落地。
Chain-of-Generation: Progressive Latent Diffusion for Text-Guided Molecular Design
- 将文本提示拆成有序步骤,逐步引导生成过程
- 在多个基准上实现更高语义对齐与多样性
- 无需训练即可提升生成透明度,适合科研设计
文本条件分子生成旨在将自然语言描述转化为化学结构,使科学家无需手工规则即可指定功能基团、骨架和理化约束。基于扩散的模型,特别是潜在扩散模型(LDM),通过在紧凑的连续潜在空间中进行随机搜索,展现出潜力。然而,现有方法依赖一次性条件编码,整个提示一次性输入并贯穿生成过程,难以满足所有要求。我们指出三个核心挑战:生成组分可解释性差、子结构遗漏、同时考虑全部要求过于激进。为此提出三原则,并设计训练无关的多阶段框架链式生成(CoG)。CoG将每个提示分解为课程顺序的语义段,逐段作为中间目标,引导去噪轨迹逐步满足更复杂的语言约束。为进一步强化语义引导,引入后对齐学习阶段,增强文本与分子潜在空间的对应关系。在基准与真实任务上的大量实验表明,CoG相比单次条件基线,在语义对齐、多样性和可控性方面表现更优,能更忠实反映复杂组合型提示,并提供生成过程的透明洞察。
原文摘要 · Abstract (English)
Text-conditioned molecular generation aims to translate natural-language descriptions into chemical structures, enabling scientists to specify functional groups, scaffolds, and physicochemical constraints without handcrafted rules. Diffusion-based models, particularly latent diffusion models (LDMs), have recently shown promise by performing stochastic search in a continuous latent space that compactly captures molecular semantics. Yet existing methods rely on one-shot conditioning, where the entire prompt is encoded once and applied throughout diffusion, making it hard to satisfy all the requirements in the prompt. We discuss three outstanding challenges of one-shot conditioning generation, including the poor interpretability of the generated components, the failure to generate all substructures, and the overambition in considering all requirements simultaneously. We then propose three principles to address those challenges, motivated by which we propose Chain-of-Generation (CoG), a training-free multi-stage latent diffusion framework. CoG decomposes each prompt into curriculum-ordered semantic segments and progressively incorporates them as intermediate goals, guiding the denoising trajectory toward molecules that satisfy increasingly rich linguistic constraints. To reinforce semantic guidance, we further introduce a post-alignment learning phase that strengthens the correspondence between textual and molecular latent spaces. Extensive experiments on benchmark and real-world tasks demonstrate that CoG yields higher semantic alignment, diversity, and controllability than one-shot baselines, producing molecules that more faithfully reflect complex, compositional prompts while offering transparent insight into the generation process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。