arXiv:2605.04709cs.LGcs.RO2026-05被引 2

ELVIS通过多假设规划与自适应误差控制,实现视觉强化学习中的长时程稳定控制。

ELVIS: Ensemble-Calibrated Latent Imagination for Long-Horizon Visual MPC

论文配图:ELVIS: Ensemble-Calibrated Latent Imagination for Long-Horizon Visual MPC
图 1 · 摘自论文原文
  • 采用高斯混合分布的MPPI,在长程规划中保持多个一致假设,避免模式平均
  • 引入不确定性感知的集成λ回报,动态调节预测深度以抑制误差累积
  • 在14个视觉任务上超越现有方法,零样本迁移至真实世界喷沙任务表现优异

基于模型的强化学习在视觉控制中面临长期规划难题:学习的潜在动力学模型在长轨迹推演中出现分支未来和多模态动作价值分布。此外,视觉遮挡导致的模型误差累积使深度想象变得脆弱。本文提出ELVIS,一种用于长时程规划的潜在模型预测控制器。该方法在Dreamer风格的循环状态空间模型(RSSM)中进行规划,并用高斯混合分布的模型预测路径积分(GMM-MPPI)替代标准的单峰方法,可在长时程中维持多个一致的假设,避免分支情况下的模式平均。同时,通过共享的不确定性感知λ回报机制——一组潜在评论家定义了上置信界(UCB)得分,动态调整时间可变的λ值,自适应平衡自举与前瞻,限制规划过程中的误差累积。该回报同时用于从想象轨迹中训练行为-评论家先验,并在GMM-MPPI中评分候选轨迹,使强化学习目标与规划器的长期优化对齐。在14个DeepMind Control Suite视觉任务上,ELVIS性能优于TD-MPC2和DreamerV3。最终,其零样本迁移到一个存在严重遮挡的真实世界喷沙任务中,显著提升表面质量指标,验证了其超越仿真环境的鲁棒性。

原文摘要 · Abstract (English)

A central challenge of visual control with model-based reinforcement learning (RL) is reliable long-horizon planning: long rollouts with learned latent dynamics exhibit branching futures and multi-modal action-value distributions. In addition, compounding model errors amplified by visual occlusions make deep imagination brittle. We present ELVIS, a latent model predictive controller (MPC) designed to make long-horizon planning practical. ELVIS plans in a Dreamer-style recurrent state space model (RSSM) and replaces standard unimodal model predictive path integral (MPPI) with a Gaussian-mixture MPPI that maintains multiple coherent hypotheses over long horizons, avoiding mode averaging under branching rollouts. In parallel, ELVIS stabilizes deep imagination with a shared uncertainty-aware lambda-return: an ensemble of latent critics defines an upper-confidence-bound (UCB) score that gates a time-varying lambda, adaptively trading off bootstrapping versus look-ahead to limit compounding error during planning. The same return is used both to train an actor-critic prior from imagined rollouts and to score candidate trajectories inside GMM-MPPI, aligning RL objectives with the planner's long-horizon optimization. On fourteen DeepMind Control Suite visual tasks, ELVIS establishes state-of-the-art performance compared with TD-MPC2 and DreamerV3. Finally, ELVIS transfers zero-shot to a real-world sand-spraying task with severe occlusions, improving surface-quality metrics and demonstrating robustness beyond simulation.

视觉控制长程规划模型预测强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。