arXiv:2505.19923cs.LGstat.ML2025-05ICML被引 5

根据数据质量动态调整正则化强度,提升离线强化学习效果

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL

  • 按状态自适应调节正则化强度,信任高质量状态的Bellman更新
  • 在D4RL上显著优于当前最佳方法,尤其在离线到在线迁移中表现突出
  • 适用于追求高精度策略学习的研究者和工业落地场景

离线强化学习旨在从静态数据集中学习有效策略。为缓解外推误差,现有方法通常对所有状态统一正则化价值函数或策略更新。然而,由于数据质量存在显著差异,固定正则化强度常导致两难:弱正则化无法解决外推误差与价值过估计,强正则化则使策略学习偏向行为克隆,抑制Bellman更新带来的性能潜力。为此,本文提出选择性状态自适应正则化方法。具体地,引入状态自适应正则化系数,以信任状态级的Bellman驱动结果,同时对高质量动作选择性施加正则化,避免低质量动作受严苛约束导致性能下降。通过建立代表值正则化方法CQL与显式策略约束方法之间的联系,将该方法有效扩展至两类主流离线强化学习框架。大量实验表明,该方法在D4RL基准上的离线及离线到在线设置中均显著优于当前最优方法。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset. To alleviate extrapolation errors, existing studies often uniformly regularize the value function or policy updates across all states. However, due to substantial variations in data quality, the fixed regularization strength often leads to a dilemma: Weak regularization strength fails to address extrapolation errors and value overestimation, while strong regularization strength shifts policy learning toward behavior cloning, impeding potential performance enabled by Bellman updates. To address this issue, we propose the selective state-adaptive regularization method for offline RL. Specifically, we introduce state-adaptive regularization coefficients to trust state-level Bellman-driven results, while selectively applying regularization on high-quality actions, aiming to avoid performance degradation caused by tight constraints on low-quality actions. By establishing a connection between the representative value regularization method, CQL, and explicit policy constraint methods, we effectively extend selective state-adaptive regularization to these two mainstream offline RL approaches. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art approaches in both offline and offline-to-online settings on the D4RL benchmark.

离线RL正则化策略优化D4RL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。