arXiv:2607.22201eess.SYcs.LG2026-07

通过KL散度约束轨迹分布,实现性能与参考轨迹的平衡控制

Trajectory-Regularized Stochastic Optimal Control via KL Divergence

论文配图:Trajectory-Regularized Stochastic Optimal Control via KL Divergence
图 1 · 摘自论文原文
  • 用KL散度正则化受控轨迹与参考轨迹分布
  • 线性二次模型下有闭式解,控制成本被扩展
  • 可调参数平衡性能优化与轨迹复现,适合数据驱动场景

我们提出轨迹正则化的随机最优控制(TRSOC),在标准随机最优控制(SOC)基础上引入受控轨迹分布与参考轨迹分布之间的Kullback-Leibler(KL)散度。利用Girsanov定理,轨迹KL散度可转化为二次漂移不匹配惩罚项,从而得到保持动态规划(DP)结构的修正运行代价。推导出相应的哈密顿-雅可比-贝尔曼(HJB)方程,并刻画了最优策略。在线性二次(LQ)设定下,该公式具有闭式解,且控制代价被扩展。实验表明,正则化参数在性能驱动与参考轨迹保持之间建立权衡,包括从离线数据学习参考动态的情况。

原文摘要 · Abstract (English)

We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions. Using Girsanov's theorem, the trajectory KL reduces to a quadratic drift mismatch penalty, yielding a modified running cost that preserves the dynamic programming (DP) structure. We derive the corresponding Hamilton--Jacobi--Bellman (HJB) equation and characterize the optimal policy. In the linear-quadratic (LQ) setting, the formulation admits a closed-form solution with an augmented control cost. Experiments show that the regularization parameter induces a trade-off between performance-driven and reference-preserving behavior, including cases with reference dynamics learned from offline data.

最优控制强化学习轨迹正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。