提出GAGA方法,让3D分子生成更快更准。
GAGA: Gaussianity-Aware Gaussian Approximation for Efficient 3D Molecular Generation
- 根据数据达到高斯性的时机,动态截断生成轨迹
- 生成质量与效率双提升,采样步数减少90%以上
- 适合需要高效生成3D分子的药物研发人员
基于高斯概率路径的生成模型(GPPGMs)通过逐步添加高斯噪声来生成数据。尽管在3D分子生成上表现优异,但其部署受限于长生成轨迹带来的高昂计算成本,训练和采样常需数百至数千步。本文提出一种名为GAGA的原理性方法,在不牺牲训练精细度或推理保真度的前提下,显著提升生成效率。核心洞察是:不同数据模态在前向过程中达到足够高斯性的时间点显著不同。据此,我们通过分析确定了分子数据获得充分高斯性的特征步数,此后可用闭式高斯近似替代原轨迹。与现有加速方法通过粗化或重构轨迹不同,GAGA保留完整分辨率的学习动力学,避免在截断分布状态间的冗余传输。在多个3D分子生成基准测试中,GAGA实现了生成质量与计算效率的显著提升。
原文摘要 · Abstract (English)
Gaussian Probability Path based Generative Models (GPPGMs) generate data by reversing a stochastic process that progressively corrupts samples with Gaussian noise. Despite state-of-the-art results in 3D molecular generation, their deployment is hindered by the high cost of long generative trajectories, often requiring hundreds to thousands of steps during training and sampling. In this work, we propose a principled method, named GAGA, to improve generation efficiency without sacrificing training granularity or inference fidelity of GPPGMs. Our key insight is that different data modalities obtain sufficient Gaussianity at markedly different steps during the forward process. Based on this observation, we analytically identify a characteristic step at which molecular data attains sufficient Gaussianity, after which the trajectory can be replaced by a closed-form Gaussian approximation. Unlike existing accelerators that coarsen or reformulate trajectories, our approach preserves full-resolution learning dynamics while avoiding redundant transport through truncated distributional states. Experiments on 3D molecular generation benchmarks demonstrate that our GAGA achieves substantial improvement on both generation quality and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。