arXiv:2505.17741cs.LGstat.ML2025-05NeurIPS被引 11

提出可训练的离散采样框架,提升采样效率与收敛速度。

Discrete Neural Flow Samplers with Locally Equivariant Transformer

  • 基于连续时间马尔可夫链建模,学习速率矩阵以满足柯尔莫哥洛夫方程。
  • 引入控制变量降低采样方差,实现稳定高效的坐标下降训练。
  • 设计局部等变Transformer,兼顾计算效率与模型表达能力,适合复杂任务。

从非归一化离散分布中采样是多个领域的基础问题。虽然马尔可夫链蒙特卡洛方法具有理论完备性,但常因混合缓慢和收敛差而受限。本文提出离散神经流采样器(DNFS),一种可训练且高效的离散采样框架。DNFS通过学习连续时间马尔可夫链的速率矩阵,使系统动态满足柯尔莫哥洛夫方程。由于目标函数涉及不可计算的分区函数,我们采用控制变量法降低其蒙特卡洛估计的方差,进而推导出坐标下降学习算法。为进一步提升计算效率,我们提出局部等变Transformer,一种新颖的速率矩阵参数化方式,在保持强表达力的同时显著加快训练速度。实验表明,DNFS在多种场景中表现优异,包括非归一化分布采样、离散能量模型训练及组合优化问题求解。

原文摘要 · Abstract (English)

Sampling from unnormalised discrete distributions is a fundamental problem across various domains. While Markov chain Monte Carlo offers a principled approach, it often suffers from slow mixing and poor convergence. In this paper, we propose Discrete Neural Flow Samplers (DNFS), a trainable and efficient framework for discrete sampling. DNFS learns the rate matrix of a continuous-time Markov chain such that the resulting dynamics satisfy the Kolmogorov equation. As this objective involves the intractable partition function, we then employ control variates to reduce the variance of its Monte Carlo estimation, leading to a coordinate descent learning algorithm. To further facilitate computational efficiency, we propose locally equivaraint Transformer, a novel parameterisation of the rate matrix that significantly improves training efficiency while preserving powerful network expressiveness. Empirically, we demonstrate the efficacy of DNFS in a wide range of applications, including sampling from unnormalised distributions, training discrete energy-based models, and solving combinatorial optimisation problems.

离散采样马尔可夫链Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。