arXiv:2508.04225cs.LGcs.AI2025-08

提出对称正则化策略优化新方法,提升离线强化学习稳定性与性能。

Symmetric Behavior Regularized Policy Optimization

  • 用无限级数近似任意f-散度,统一处理对称与非对称正则化。
  • 推导出闭式最优策略表达式,实现数值稳定优化。
  • 在D4RL基准上表现稳健,对近似项数不敏感,适合实际部署。

行为正则化策略优化(BRPO)通过非对称散度正则化缓解离线强化学习中的分布偏移问题。本文首次探讨对称BRPO这一开放问题。通过教学案例表明,对称正则化在应对单边偏差、边界附近策略更新及投影几何一致性方面优于非对称正则化。然而,对称散度在BRPO中不自然:无法获得闭式解,且作为优化目标时易引发数值不稳定性。为此,我们提出一个通用的BRPO框架,采用皮尔逊-瓦贾达散度的无穷级数表示任意f-散度,涵盖对称与非对称情形。通过有限项近似,得到以下结果:(1) 闭式最优策略表达式;(2) 数值稳定的优化代理;(3) 近似精度的紧上界。在D4RL基准和教学案例中,所提方法表现持续优异,且对近似项数不敏感。

原文摘要 · Abstract (English)

Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline reinforcement learning. This paper is the first to study the open question of symmetric BRPO. Using didactic examples, we show that symmetric regularization can outperform asymmetric regularization in addressing one-sided bias, near-boundary policy updates, and projection geometry consistency. However, symmetric divergences do not fit BRPO naturally: they do not permit a closed-form solution when used as regularizers, and can lead to numerical instability when used as optimization objectives. We first introduce a universal BRPO framework using an infinite series of Pearson-Vajda divergences to represent any $f$-divergence, which includes both symmetric and asymmetric divergences. We use a finite-series approximation to obtain the following results for symmetric BRPO: (1) a closed-form optimal policy expression; (2) a numerically stable optimization surrogate; and (3) a tight upper bound on the approximation quality. On the D4RL benchmark and in didactic examples, we show that the proposed method achieves consistently strong results and is robust to the number of terms in the approximation.

强化学习离线学习策略优化正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。