arXiv:2506.08644cs.LG2025-06

改进离线约束强化学习中的成本估计,提升评估与优化准确性。

Semi-gradient DICE for Offline Constrained Reinforcement Learning

  • 提出半梯度DICE方法,解决现有算法在离线评估中的偏差问题。
  • 在DSRL基准上实现当前最优性能,成本估计误差显著降低。
  • 适合需要可靠成本评估的离线约束强化学习任务。

静态分布修正估计(DICE)解决了策略诱导的静态分布与离线策略评估(OPE)和策略优化所需目标分布之间的不匹配问题。基于DICE的离线约束强化学习尤其受益于其灵活性,可在离线设置中同时最大化回报并估计成本。然而我们发现,近期为提升DICE框架离线强化学习性能的方法无意中削弱了其进行OPE的能力,因此不适用于约束强化学习场景。本文揭示了该限制的根本原因:依赖半梯度优化,导致求解的是一个根本不同的优化问题,从而造成成本估计失败。基于此洞察,我们提出一种新方法,使半梯度DICE能够实现可靠的OPE和约束强化学习。所提方法确保了准确的成本估计,并在离线约束强化学习基准DSRL上达到当前最优表现。

原文摘要 · Abstract (English)

Stationary Distribution Correction Estimation (DICE) addresses the mismatch between the stationary distribution induced by a policy and the target distribution required for reliable off-policy evaluation (OPE) and policy optimization. DICE-based offline constrained RL particularly benefits from the flexibility of DICE, as it simultaneously maximizes return while estimating costs in offline settings. However, we have observed that recent approaches designed to enhance the offline RL performance of the DICE framework inadvertently undermine its ability to perform OPE, making them unsuitable for constrained RL scenarios. In this paper, we identify the root cause of this limitation: their reliance on a semi-gradient optimization, which solves a fundamentally different optimization problem and results in failures in cost estimation. Building on these insights, we propose a novel method to enable OPE and constrained RL through semi-gradient DICE. Our method ensures accurate cost estimation and achieves state-of-the-art performance on the offline constrained RL benchmark, DSRL.

强化学习离线学习成本估计约束控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。