用扩散桥建模价值分布,提升强化学习的评估精度。
Distributional Reinforcement Learning with Diffusion Bridge Critics
- 用扩散桥直接建模价值的逆累积分布函数,避免分布坍缩。
- 在MuJoCo上优于已有分布式评论者模型,尤其在复杂任务中表现更稳。
- 可即插即用,适配主流强化学习框架,适合追求高精度评估的研究者。
基于扩散的强化学习方法在连续控制任务中展现出良好前景,但现有研究主要关注扩散策略,未探索扩散评论者。由于策略优化依赖评论者,准确的价值估计比策略表达能力更为关键。考虑到多数强化学习任务具有随机性,分布式建模更适合评论者。为此,本文提出分布式强化学习新方法——扩散桥评论者(DBC),直接建模Q值的逆累积分布函数。该方法能精确捕捉价值分布,凭借扩散桥强大的分布匹配能力,防止其坍缩为平凡高斯分布。此外,我们推导出解析积分公式以解决离散化误差,这对价值估计至关重要。据我们所知,DBC是首个将扩散桥用于评论者的成果。值得注意的是,DBC可作为即插即用模块集成到多数现有强化学习框架中。在MuJoCo机器人控制基准上的实验表明,DBC显著优于以往分布式评论者模型。
原文摘要 · Abstract (English)
Recent advances in diffusion-based reinforcement learning (RL) methods have demonstrated promising results in a wide range of continuous control tasks. However, existing works in this field focus on the application of diffusion policies while leaving the diffusion critics unexplored. In fact, since policy optimization fundamentally relies on the critic, accurate value estimation is far more important than policy expressiveness. Furthermore, given the stochasticity of most reinforcement learning tasks, it has been confirmed that the critic is more appropriately depicted with a distributional model. Motivated by these points, we propose a novel distributional RL method with Diffusion Bridge Critics (DBC). DBC directly models the inverse cumulative distribution function (CDF) of the Q value. This allows us to accurately capture the value distribution and prevents it from collapsing into a trivial Gaussian distribution owing to the strong distribution-matching capability of the diffusion bridge. Moreover, we further derive an analytic integral formula to address discretization errors in DBC, which is essential in value estimation. To our knowledge, DBC is the first work to employ the diffusion bridge model as the critic. Notably, DBC is also a plug-and-play component and can be integrated into most existing RL frameworks. Experimental results on MuJoCo robot control benchmarks demonstrate the superiority of DBC compared with previous distributional critic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。