统一了最优控制中的熵正则化,打通经典与现代方法的联系。
Unifying Entropy Regularization in Optimal Control: From and Back to Classical Objectives via Iterated Soft Policies and Path Integral Solutions
- 分离策略与转移的KL正则项,用独立权重实现统一建模
- 软策略形式可逼近原始目标,迭代求解能恢复经典控制问题
- 特定情形下获得路径积分解与可组合性,计算更高效
本文通过Kullback-Leibler(KL)正则化的视角,构建了一个统一的最优控制框架。提出一个核心问题,将策略和转移过程的KL惩罚项分开处理,并赋予独立权重,从而推广了传统轨迹级的KL正则化方法。该统一框架可还原多种控制问题:经典随机最优控制(SOC)、风险敏感随机最优控制(RSOC),以及对应的基于策略的KL正则化版本——软策略SOC和软策略RSOC,这些版本提供了可计算的替代目标。值得注意的是,这些软策略形式对原问题具有上界性质,因此迭代求解可逐步收敛至原始目标。此外,在软策略RSOC中发现一种同步情形:当策略与转移的KL权重相等时,可导出线性贝尔曼算子、路径积分解及组合性,使这类计算优势扩展到更广泛的控制问题。
原文摘要 · Abstract (English)
This paper develops a unified perspective on several optimal control formulations through the lens of Kullback-Leibler (KL) regularization. We propose a central problem that separates the KL penalties on policies and transitions with independent weights, thus generalizing the standard trajectory-level KL-regularization used in probabilistic optimal control. This umbrella formulation recovers various control problems: the classical Stochastic Optimal Control (SOC), Risk-Sensitive Stochastic Optimal Control (RSOC), and their policy-based KL-regularized counterparts, termed soft-policy SOC and RSOC, which yield tractable surrogates. Beyond being regularized variants, these soft-policy formulations majorize the original SOC and RSOC, thus, iterating their solutions recovers the original objectives. We further identify a synchronized case of soft-policy RSOC where the policy and transition KL weights coincide, yielding a linear Bellman operator, path-integral solution, and compositionality -- extending these computationally favourable properties to a broad class of control problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。