arXiv:2607.14245eess.SYcs.RO2026-07

用熵反馈自动调温,让确定性MPPI更快更稳地找到最优解。

Information-Theoretic Adaptive Cooling for Deterministic MPPI via Entropy Feedback

论文配图:Information-Theoretic Adaptive Cooling for Deterministic MPPI via Entropy Feedback
图 1 · 摘自论文原文
  • 根据重要性权重的香农熵动态调节温度,实现自适应冷却。
  • 在非光滑任务中收敛速度比现有方法快,且不依赖梯度信息。
  • 适合复杂系统运动规划,尤其适用于无梯度优化场景。

本文研究基于模型预测路径积分(MPPI)的确定性最优控制,这是一种采样驱动、无需梯度的框架,适用于具有复杂动力学和非光滑目标的系统。在确定性MPPI中,温度需降至零以逼近真实最优解,但有效的降温策略设计仍是根本挑战。现有方法通常依赖预设的开环降温方案,限制了算法效率与鲁棒性。为此,本文提出信息论自适应冷却(ITAC)框架,利用重要性权重的香农熵作为在线反馈信号调节温度。该机制根据当前采样状态自适应调整降温速率:当权重分布分散时快速下降,集中时则缓慢冷却。我们证明了该方案渐近收敛至确定性最优解,并推导出一个关键熵阈值,可平滑防止权重过早坍缩。在非光滑信号时序逻辑运动规划任务上的实验表明,ITAC显著提升采样效率,收敛速度远超现有最优基线,同时保持MPPI的无梯度特性。

原文摘要 · Abstract (English)

This paper investigates deterministic optimal control using Model Predictive Path Integral (MPPI) control, a sampling-based and derivative-free framework well suited for systems with complex dynamics and nonsmooth objectives. In deterministic MPPI, the temperature must be driven to zero to recover the true optimum, yet the design of an effective cooling schedule remains a fundamental challenge. Existing methods typically rely on predefined open-loop schedules, which limit the efficiency and robustness of the algorithm. To overcome this limitation, we propose an Information-Theoretic Adaptive Cooling (ITAC) framework that uses the Shannon entropy of the importance weights as an online feedback signal to regulate the temperature. The proposed mechanism adapts the cooling rate to the current sampling state, enabling fast progress when the weights are diffuse and cautious cooling when they become concentrated. We prove asymptotic convergence of the resulting scheme to the deterministic optimum, and further derive a critical entropy threshold that leads to a smooth barrier against premature weight collapse. Experiments on nonsmooth signal temporal logic motion-planning tasks show that ITAC improves sampling efficiency and achieves substantially faster convergence than state-of-the-art baselines without sacrificing the derivative-free nature of MPPI.

最优控制MPPI自适应冷却强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。