用新方法解决强化学习中KL正则化失效问题,让控制更稳定。
Well-Posed KL-Regularized Control via Wasserstein and Kalman-Wasserstein KL Divergences
- 用运输几何替代传统信息几何,构造新型KL散度
- 在零噪声下仍保持有限值,避免控制问题退化
- 适合研究强化学习与最优控制的学者参考
KL散度正则化广泛用于强化学习,但在支持集不匹配时会发散,且在低噪声情况下会退化。本文通过统一的信息几何框架,将动态形式中的Fisher-Rao几何替换为基于传输的几何,推导出常见分布族的闭式表达式。在椭圆分布间,这些散度在协方差相等时仍保持有限,并为卡尔曼集合方法中的正则化启发式提供了几何解释。我们在KL正则化最优控制中验证了其有效性。在线性时不变系统与高斯过程噪声的可解析设置下,经典KL退化为奇异的二次控制惩罚;而本文提出的方法消除了这种奇异性,使问题具有良定性。在双积分器和倒立摆实例中,所得控制保持非平凡反馈,且闭环性能更优。
原文摘要 · Abstract (English)
Kullback-Leibler (KL) divergence regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise regimes. Using a unified information-geometric framework, we introduce KL analogs by replacing the Fisher-Rao geometry in the dynamical formulation of the KL with transport-based geometries, and derive closed-form expressions for common distribution families. Between elliptic distributions, these divergences remain finite for degenerating equal covariances and yield a geometric interpretation of regularization heuristics used in Kalman ensemble methods. We demonstrate the utility of these divergences in KL-regularized optimal control. In the fully tractable setting of linear time-invariant systems with Gaussian process noise, the classical KL reduces to a quadratic control penalty that becomes singular as process noise vanishes. Our variants remove this singularity and yield well-posed problems. In both the double integrator and cart-pole examples, the resulting controls preserve nontrivial feedback and achieve better closed-loop performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。