通过稳定价值函数梯度场,提升强化学习策略平滑性。
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods
- 从评论家视角出发,用梯度场稳定性约束策略平滑性。
- 实验证明该方法在不改策略网络下实现与传统方法相当的平滑性。
- 适合需要稳定策略输出的物理部署场景,如机器人控制。
连续动作演员-评论家方法学到的策略常出现高频振荡,难以用于真实物理系统。现有方法通过直接正则化策略输出来强制平滑,但治标不治本。本文理论证明:策略非平滑性本质上由评论家的微分几何特性决定。通过对演员-评论家目标应用隐式微分,我们推导出最优策略敏感度受Q函数混合偏导(噪声敏感性)与动作空间曲率(信号区分度)之比限制。为验证该理论,提出基于评论家视角的PAVE框架,将评论家视为标量场,稳定其诱导的动作梯度场,在抑制Q梯度波动的同时保留局部曲率。实验表明,PAVE在不修改演员网络的前提下,达到与策略侧正则化相当的平滑性,且任务性能保持竞争力。
原文摘要 · Abstract (English)
Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physical deployment. Current approaches attempt to enforce smoothness by directly regularizing the policy's output. We argue that this approach treats the symptom rather than the cause. In this work, we theoretically establish that policy non-smoothness is fundamentally governed by the differential geometry of the critic. By applying implicit differentiation to the actor-critic objective, we prove that the sensitivity of the optimal policy is bounded by the ratio of the Q-function's mixed-partial derivative (noise sensitivity) to its action-space curvature (signal distinctness). To empirically validate this theoretical insight, we introduce PAVE (Policy-Aware Value-field Equalization), a critic-centric regularization framework that treats the critic as a scalar field and stabilizes its induced action-gradient field. PAVE rectifies the learning signal by minimizing the Q-gradient volatility while preserving local curvature. Experimental results demonstrate that PAVE achieves smoothness comparable to policy-side smoothness regularization methods, while maintaining competitive task performance, without modifying the actor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。