arXiv:2608.08644cs.LG2026-08

提出路径依赖的离散推断方法,提升采样效率与探索能力。

Path-dependent Discrete Amortized Inference

论文配图:Path-dependent Discrete Amortized Inference
图 1 · 摘自论文原文
  • 用可学习的动态系统替代马尔可夫状态,让策略依赖完整历史轨迹。
  • 在标准任务上收敛更快,状态空间探索能力显著优于现有方法。
  • 适合需要高效生成离散组合对象的任务,如结构化推理与程序合成。

我们研究从给定的非归一化后验分布中采样组合型离散对象的问题。近期研究表明,可通过学习一个确定性马尔可夫决策过程(MDP)来逐步构建符合后验的对象,实现高效采样。然而,本文指出马尔可夫假设会阻碍训练中的信号传播,并因状态混淆导致学习到的采样器表达能力严重下降。为此,我们提出将MDP提升为可学习的潜在动态系统,使底层策略依赖于完整的历史路径而非仅当前状态。因此,我们将该方法称为路径依赖的离散摊销推断。更重要的是,我们证明可将现有的离散摊销采样器学习算法推广至本设定。在标准基准任务上的实验表明,我们的方法通常能实现更快的学习收敛和更优的状态空间探索能力。

原文摘要 · Abstract (English)

We consider the problem of sampling compositional and discrete objects from a given unnormalized posterior distribution. Notably, recent studies have shown that this problem can be efficiently solved by learning a deterministic Markov Decision Process (MDP) that progressively builds each object in proportion to the posterior. In this work, however, we demonstrate that the Markovian assumption can both hamper signal propagation during training and catastrophically reduce the learned sampler's expressivity due to state aliasing. To address these issues, we propose lifting the MDP with a learnable latent dynamical system that allows the underlying policy to depend on the entire past trajectory---and not only on the current state. In view of this, we refer to the resulting method as path-dependent discrete amortized inference. Importantly, we provably extend existing learning algorithms for discrete amortized samplers to our setting. In experiments on standard benchmark problems, we also show that our approach often leads to faster learning convergence and improved state space exploration relatively to prior techniques.

离散采样路径依赖动态系统摊销推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。