arXiv:2605.07520cs.AI2026-05

通过动态噪声提升复杂系统决策优化效果

Model-Driven Policy Optimization in Differentiable Simulators via Stochastic Exploration

论文配图:Model-Driven Policy Optimization in Differentiable Simulators via Stochastic Exploration
图 1 · 摘自论文原文
  • 在可微仿真中注入随梯度变化的随机动作噪声
  • 在非线性混合域中显著提升解的质量,优于现有方法
  • 适合需要高精度决策的强化学习与规划场景

可微规划通过系统动力学的可微模型实现基于梯度的决策优化。然而,在高度非线性和混合离散-连续域中,优化景观常存在平坦区域与突变边界,阻碍有效优化。我们提出模型驱动策略优化(MDPO),通过在优化过程中向动作空间注入噪声,引入随机探索。利用对模型的访问,MDPO根据轨迹目标的梯度敏感性自适应调整噪声幅度,形成随时间变化的探索策略。这提升了目标景观的探索能力,并通过跨时间步与迭代动态分配探索,帮助跳出劣质局部最优。基准测试表明,MDPO在挑战性的非线性与混合设置中,持续优于无噪声的确定性可微规划及当前最先进实现,也优于如PPO等模型无关基线,显著提升解质量。我们进一步分析了自适应噪声幅度在时间步与优化迭代中的演化过程,揭示了探索分配机制。

原文摘要 · Abstract (English)

Differentiable planning enables gradient-based optimization of decision-making problems by leveraging differentiable models of system dynamics. However, in highly nonlinear and hybrid discrete-continuous domains, the resulting optimization landscapes are often ill-conditioned, with flat regions and sharp transitions that hinder effective optimization. We propose Model-Driven Policy Optimization (MDPO), a framework that introduces stochastic exploration into differentiable planning by injecting noise into the action space during optimization. Leveraging access to the model, MDPO further adapts the noise magnitude based on gradient-derived sensitivity of the trajectory objective, yielding a time-dependent exploration profile. This enables improved exploration of the objective landscape and helps escape poor local optima via dynamic allocation of exploration across timesteps and iterations. Experiments on benchmark domains demonstrate that MDPO consistently outperforms deterministic differentiable planning, including both the noise-free variant of our method and available state-of-the-art implementations, as well as model-free baselines such as PPO, significantly improving solution quality across challenging nonlinear and hybrid settings. We further analyze the evolution of the adaptive noise magnitude across both time steps and optimization iterations, providing insight into how exploration is allocated during learning.

强化学习可微规划策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。