提出新型扩散框架,让分子生成与描述更精准。
BiMol-Diff: A Unified Diffusion Framework for Molecular Generation and Captioning

- 根据分子片段恢复难度动态调整噪声,保护关键结构。
- 在ChEBI-20和M3-20M数据集上,精确匹配率提升15.4%。
- 适合需要高精度分子语言建模的研究者使用。
连接分子结构与自然语言对可控设计至关重要。自回归模型难以处理长程依赖,而标准扩散过程对所有位置施加相同噪声,会破坏具有结构信息的标记。我们提出BiMol-Diff,一个统一的扩散框架,用于文本条件下的分子生成与分子描述任务。核心是基于标记恢复难度的位置感知噪声调度,在前向过程中保留更难恢复的子结构。在ChEBI-20和M3-20M数据集上,该方法在分子重建上实现15.4%的相对精度提升,并在描述生成任务中达到最优的BLEU和BERTScore表现。结果表明,标记感知的噪声机制能显著提升分子结构-语言建模的保真度。
原文摘要 · Abstract (English)
Bridging molecular structures and natural language is essential for controllable design. Autoregressive models struggle with long-range dependencies, while standard diffusion processes apply uniform corruption across positions, which can distort structurally informative tokens. We present BiMol-Diff, a unified diffusion framework for the paired tasks of text-conditioned molecule generation and molecule captioning. Our key component is a token-aware noise schedule that assigns position-dependent corruption based on token recovery difficulty, preserving harder-to-recover substructures during the forward process. On ChEBI-20 and M3-20M, BiMol-Diff improves molecule reconstruction with a 15.4% relative gain in Exact Match and achieves strong captioning results, attaining best BLEU and BERTScore among compared baselines. These results indicate token-aware noising improves fidelity in molecular structure-language modelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。