arXiv:2410.01796cs.LG2024-10被引 3

让生成模型在线性空间中工作,解决强化学习中的分布建模难题

Bellman Diffusion: Generative Modeling as Learning a Linear Operator in the Distribution Space

  • 用梯度与标量场建模保持马尔可夫决策过程的线性特性
  • 在分布强化学习任务中收敛速度比传统方法快1.5倍
  • 适合需要高精度分布建模的强化学习与生成任务

深度生成模型(DGMs),包括基于能量的模型(EBMs)和基于得分的生成模型(SGMs),在高保真数据生成和复杂连续分布逼近方面取得了进展。然而,它们在马尔可夫决策过程(MDPs)中的应用,尤其是在分布强化学习(RL)领域,仍处于探索阶段,传统直方图方法仍占主导地位。本文严谨指出,这一应用差距源于现代DGM的非线性特性与MDPs中贝尔曼方程所需的线性要求相冲突。例如,EBMs涉及能量函数的指数运算和归一化常数等非线性操作。为解决此问题,我们提出贝尔曼扩散(Bellman Diffusion),一种新的DGM框架,通过梯度与标量场建模维持MDPs中的线性性。结合基于散度的训练技术优化神经网络代理,并引入新型随机微分方程(SDE)进行采样,贝尔曼扩散保证收敛到目标分布。实验表明,该方法能实现精确的场估计,且具备良好的图像生成能力,在分布强化学习任务中收敛速度比传统直方图基线快1.5倍。本工作实现了DGM在MDP应用中的有效整合,为先进决策框架开辟新路径。

原文摘要 · Abstract (English)

Deep Generative Models (DGMs), including Energy-Based Models (EBMs) and Score-based Generative Models (SGMs), have advanced high-fidelity data generation and complex continuous distribution approximation. However, their application in Markov Decision Processes (MDPs), particularly in distributional Reinforcement Learning (RL), remains underexplored, with conventional histogram-based methods dominating the field. This paper rigorously highlights that this application gap is caused by the nonlinearity of modern DGMs, which conflicts with the linearity required by the Bellman equation in MDPs. For instance, EBMs involve nonlinear operations such as exponentiating energy functions and normalizing constants. To address this, we introduce Bellman Diffusion, a novel DGM framework that maintains linearity in MDPs through gradient and scalar field modeling. With divergence-based training techniques to optimize neural network proxies and a new type of stochastic differential equation (SDE) for sampling, Bellman Diffusion is guaranteed to converge to the target distribution. Our empirical results show that Bellman Diffusion achieves accurate field estimations and is a capable image generator, converging 1.5x faster than the traditional histogram-based baseline in distributional RL tasks. This work enables the effective integration of DGMs into MDP applications, unlocking new avenues for advanced decision-making frameworks.

生成模型强化学习分布建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。