arXiv:2602.05323cs.LGcs.AI2026-02

用生成模型提升离线安全强化学习的奖励与约束平衡能力

GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL

  • 通过重标注和增强数据,实现从次优轨迹中拼接优质动作序列
  • 利用期望回归估计最优奖赏与成本目标,实现更优的权衡效果
  • 适合需要在约束下高效训练策略的研究者和工程应用

离线安全强化学习(OSRL)旨在仅使用预收集数据集学习高绩效策略并满足约束条件。受生成模型(GMs)能力启发,近期方法将决策过程重构为条件生成任务,由GM生成给定奖赏与成本值下的理想动作。然而,现有方法面临两大挑战:(1)难以从数据集中“拼接”出最优转移路径;(2)在奖赏与成本目标冲突时难以平衡。为此,本文提出目标辅助拼接(GAS)算法,通过在转移层面上增强并重标注数据,构建高质量轨迹。GAS引入新型目标函数,基于重标注数据上的期望回归,估计数据集中可实现的最优奖赏与成本目标,从而支持更广泛的奖赏-成本回报组合,并优于人工设定值。这些估计目标引导策略训练,在约束环境下实现稳健性能。此外,通过重塑数据分布使奖赏-成本回报更均匀,提升训练稳定性和效率。实验验证了GAS在平衡奖赏最大化与约束满足方面的优越性。

原文摘要 · Abstract (English)

Offline Safe Reinforcement Learning (OSRL) aims to learn a policy to achieve high performance in sequential decision-making while satisfying constraints, using only pre-collected datasets. Recent works, inspired by the strong capabilities of Generative Models (GMs), reformulate decision-making in OSRL as a conditional generative process, where GMs generate desirable actions conditioned on predefined reward and cost values. However, GM-assisted methods face two major challenges in OSRL: (1) lacking the ability to "stitch" optimal transitions from suboptimal trajectories within the dataset, and (2) struggling to balance reward targets with cost targets, particularly when they are conflict. To address these issues, we propose Goal-Assisted Stitching (GAS), a novel algorithm designed to enhance stitching capabilities while effectively balancing reward maximization and constraint satisfaction. To enhance the stitching ability, GAS first augments and relabels the dataset at the transition level, enabling the construction of high-quality trajectories from suboptimal ones. GAS also introduces novel goal functions, which estimate the optimal achievable reward and cost goals from the dataset. These goal functions, trained using expectile regression on the relabeled and augmented dataset, allow GAS to accommodate a broader range of reward-cost return pairs and achieve a better tradeoff between reward maximization and constraint satisfaction compared to human-specified values. The estimated goals then guide policy training, ensuring robust performance under constrained settings. Furthermore, to improve training stability and efficiency, we reshape the dataset to achieve a more uniform reward-cost return distribution. Empirical results validate the effectiveness of GAS, demonstrating superior performance in balancing reward maximization and constraint satisfaction compared to existing methods.

强化学习生成模型安全控制离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。