arXiv:2410.12522cs.LG2024-10

用函数空间建模分子生成,速度更快且更准确

MING: A Functional Approach to Learning Molecular Generative Models

  • 在函数空间而非数据空间进行扩散生成,简化模型结构
  • 在多个分子数据集上超越现有方法,生成速度显著提升
  • 适合需要高效生成高质量分子的研究者使用

传统分子生成方法多基于序列或图结构表示,限制表达能力或需复杂对称架构。本文提出基于函数表示的新范式——分子隐式神经生成(MING),一种基于扩散的模型,在函数空间学习分子分布。与数据空间的标准扩散过程不同,MING采用新颖的函数去噪概率过程,通过期望最大化算法联合去噪函数输入与输出空间的信息,利用数据的隐式神经表示。该方法实现简单而高效的模型设计,准确捕捉底层函数分布。在多个分子相关数据集上的实验表明,MING性能优于当前最优的数据空间方法,生成样本更合理,且架构更简洁、生成速度显著更快。代码已公开于 https://github.com/v18nguye/MING。

原文摘要 · Abstract (English)

Traditional molecule generation methods often rely on sequence- or graph-based representations, which can limit their expressive power or require complex permutation-equivariant architectures. This paper introduces a novel paradigm for learning molecule generative models based on functional representations. Specifically, we propose Molecular Implicit Neural Generation (MING), a diffusion-based model that learns molecular distributions in the function space. Unlike standard diffusion processes in the data space, MING employs a novel functional denoising probabilistic process, which jointly denoises information in both the function's input and output spaces by leveraging an expectation-maximization procedure for latent implicit neural representations of data. This approach enables a simple yet effective model design that accurately captures underlying function distributions. Experimental results on molecule-related datasets demonstrate MING's superior performance and ability to generate plausible molecular samples, surpassing state-of-the-art data-space methods while offering a more streamlined architecture and significantly faster generation times. The code is available at https://github.com/v18nguye/MING.

分子生成扩散模型函数表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。