arXiv:2503.05979cs.LGcs.AI2025-03ICML被引 32

让模型自己学生成顺序,提升分子图生成效果

Learning-Order Autoregressive Models with Application to Molecular Graph Generation

  • 用可学习的顺序策略动态决定生成步骤
  • 在QM9和ZINC250k上达到当前最优性能
  • 适合需要高质量分子结构生成的研究者

自回归模型(ARM)已成为序列生成任务的核心方法,许多问题可建模为下一步词预测。虽然文本有自然的左右顺序,但对图等高维数据,标准顺序不明确。为此,我们提出一种新型自回归模型,通过从数据中逐序推断出概率性生成顺序来生成高维数据。该模型引入一个可训练的概率分布,称为顺序策略(order-policy),根据当前状态动态决定自回归顺序。为训练模型,我们提出了对数似然的变分下界,并使用随机梯度估计进行优化。实验表明,该方法能在图像和图生成中学习到有意义的生成顺序。在分子图生成这一挑战性任务上,我们在QM9和ZINC250k基准上取得当前最优结果,各项评估指标(分布相似性、类药性)均表现优异。

原文摘要 · Abstract (English)

Autoregressive models (ARMs) have become the workhorse for sequence generation tasks, since many problems can be modeled as next-token prediction. While there appears to be a natural ordering for text (i.e., left-to-right), for many data types, such as graphs, the canonical ordering is less obvious. To address this problem, we introduce a variant of ARM that generates high-dimensional data using a probabilistic ordering that is sequentially inferred from data. This model incorporates a trainable probability distribution, referred to as an order-policy, that dynamically decides the autoregressive order in a state-dependent manner. To train the model, we introduce a variational lower bound on the log-likelihood, which we optimize with stochastic gradient estimation. We demonstrate experimentally that our method can learn meaningful autoregressive orderings in image and graph generation. On the challenging domain of molecular graph generation, we achieve state-of-the-art results on the QM9 and ZINC250k benchmarks, evaluated across key metrics for distribution similarity and drug-likeless.

分子生成自回归图生成顺序学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。