arXiv:2606.13169cs.RO2026-06

改进正则化设计,让强化学习策略更平滑且稳定。

Redesigning Regularization for Effective Policy Smoothing

  • 重新设计正则化,解决局部光滑性与表达力的权衡问题。
  • 在多任务和算法上实现平滑运动,提升控制性能。
  • 适用于四足机器人模拟到现实的迁移,增强对速度突变的鲁棒性。

本文提出一种新型正则化设计,有效实现强化学习中策略函数的平滑。尽管最初考虑通过增强全局Lipschitz连续性来实现平滑,但因平滑性与表达力之间的权衡,实际仅能实现局部Lipschitz连续性。然而,原始实现复杂且平滑效果不足,导致研究者更倾向使用简单方法。这反映出理论与实现间的脱节。本文识别出原始实现失效的三个原因,并提供相应修正方案。修改后的正则化在多个任务与算法中表现优异,成功实现平滑运动并提升控制性能。进一步应用于四足机器人模拟到现实的强化学习,证明平滑运动可有效应对目标速度命令的突发变化,增强系统鲁棒性。

原文摘要 · Abstract (English)

This paper proposes a novel regularization design to effectively smooth policy functions in reinforcement learning. While regularization that enhances ``global'' Lipschitz continuity was initially considered, it has been limited to ``local'' Lipschitz continuity due to a tradeoff between smoothness and expressiveness. However, it has become apparent that the original implementation is cumbersome and does not provide sufficient smoothing, leading to a preference for simpler implementations. This stems from a discrepancy between theory and implementation, and a more appropriate implementation can expect to facilitate smoothing. Therefore, this paper identifies three reasons why the original implementation does not function adequately and provide remedies for them. This modified regularization performs well across multiple tasks and algorithms, successfully achieving smooth motion while improving control performance. Furthermore, by applying it to sim-to-real reinforcement learning for a quadruped robot, it is demonstrated that smooth motion provides robustness against sudden changes in target velocity commands.

强化学习策略平滑正则化四足机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。