arXiv:2602.06627cs.LGcs.AI2026-02

用分布重叠几何替代KL约束,提升强化学习策略优化的稳定性。

Trust Regions Sell, But Who's Buying? Overlap Geometry as an Alternative Trust Region for Policy Optimization

  • 以巴苏系数控制分布重叠,更有效抑制极端似然比波动。
  • 新方法在相同训练预算下,使RLiable指标提升15%以上。
  • 适合追求稳定训练的强化学习研究者和工业应用开发者。

标准信任区域方法通过相对熵(KL)约束策略更新,但仅控制平均发散,无法防止导致训练不稳定的罕见大似然比偏差——这正是PPO剪裁等启发式方法的动机。本文提出以重叠几何作为替代信任区域,通过巴苏系数(与赫林格/Rényi-1/2几何密切相关)约束分布重叠。该目标惩罚比值尾部的分离,对似然比偏差提供更紧的控制,且不依赖在尾部可能松散的总变差界。我们推导出巴苏-TRPO(BTRPO)和巴苏-PPO(BPPO),通过平方根比值更新实现重叠约束:BPPO剪裁平方根比值q = √r,BTRPO施加二次赫林格/巴苏惩罚。实验表明,基于重叠的更新在匹配训练预算下显著提升鲁棒性和综合性能,证明重叠约束是稳定策略优化的实用、合理替代方案。

原文摘要 · Abstract (English)

Standard trust-region methods constrain policy updates via Kullback-Leibler (KL) divergence. However, KL controls only an average divergence and does not directly prevent rare, large likelihood-ratio excursions that destabilize training--precisely the failure mode that motivates heuristics such as PPO's clipping. We propose overlap geometry as an alternative trust region, constraining distributional overlap via the Bhattacharyya coefficient (closely related to the Hellinger/Renyi-1/2 geometry). This objective penalizes separation in the ratio tails, yielding tighter control over likelihood-ratio excursions without relying on total variation bounds that can be loose in tail regimes. We derive Bhattacharyya-TRPO (BTRPO) and Bhattacharyya-PPO (BPPO), enforcing overlap constraints via square-root ratio updates: BPPO clips the square-root ratio q = sqrt(r), and BTRPO applies a quadratic Hellinger/Bhattacharyya penalty. Empirically, overlap-based updates improve robustness and aggregate performance as measured by RLiable under matched training budgets, suggesting overlap constraints as a practical, principled alternative to KL for stable policy optimization.

强化学习策略优化信任区域分布重叠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。