arXiv:2604.21456cs.LGcs.RO2026-04

用采样方法优化轨迹与策略,通过温度退火提升效率。

Tempered Sequential Monte Carlo for Trajectory and Policy Optimization with Differentiable Dynamics

论文配图:Tempered Sequential Monte Carlo for Trajectory and Policy Optimization with Differentiable Dynamics
图 1 · 摘自论文原文
  • 将控制器设计转化为推断问题,用变分方法优化
  • 在多个基准上优于现有最先进方法,收敛更快
  • 适合需要精确梯度的强化学习与控制场景

我们提出一种基于采样的有限时域轨迹与策略优化框架,适用于可微动力学系统。将控制器设计视为推断问题,最小化带KL正则的期望轨迹成本,得到随温度降低而集中于低代价解的“玻尔兹曼加权”控制器参数分布。为高效采样这一尖锐且可能多模态的目标分布,引入温控序贯蒙特卡洛(TSMC):通过从先验到目标分布的退火路径,自适应重加权并重采样粒子,结合哈密顿蒙特卡洛再生维持多样性,并利用通过轨迹回放微分获得的精确梯度。针对策略优化,通过(i)初始状态分布的确定性经验近似和(ii)将回放随机性作为辅助变量的扩展空间构造进行拓展。在轨迹与策略优化基准测试中,TSMC展现出广泛适用性,性能优于现有最优基线。

原文摘要 · Abstract (English)

We propose a sampling-based framework for finite-horizon trajectory and policy optimization under differentiable dynamics by casting controller design as inference. Specifically, we minimize a KL-regularized expected trajectory cost, which yields an optimal "Boltzmann-tilted" distribution over controller parameters that concentrates on low-cost solutions as temperature decreases. To sample efficiently from this sharp, potentially multimodal target, we introduce tempered sequential Monte Carlo (TSMC): an annealing scheme that adaptively reweights and resamples particles along a tempering path from a prior to the target distribution, while using Hamiltonian Monte Carlo rejuvenation to maintain diversity and exploit exact gradients obtained by differentiating through trajectory rollouts. For policy optimization, we extend TSMC via (i) a deterministic empirical approximation of the initial-state distribution and (ii) an extended-space construction that treats rollout randomness as auxiliary variables. Experiments across trajectory- and policy-optimization benchmarks show that TSMC is broadly applicable and compares favorably to state-of-the-art baselines.

强化学习采样优化可微动力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。