arXiv:2602.11075cs.RO2026-02被引 34

用想象中的世界模型让机器人自我优化,提升复杂操作的可靠性。

RISE: Self-Improving Robot Policy with Compositional World Model

  • 构建可组合的世界模型,通过想象预测多视角未来并评估进展
  • 在真实任务中实现超35%性能提升,动态积木排序等任务表现显著
  • 无需真实交互即可自进化,适合需要高安全性和低试错成本场景

尽管模型规模和数据获取持续增长,视觉-语言-动作(VLA)模型在涉及接触和动态操作的任务中仍显脆弱,微小执行偏差会累积成失败。虽然强化学习(RL)提供了一条通向鲁棒性的路径,但物理世界中的在线策略强化学习受限于安全风险、硬件成本和环境重置。为此,我们提出RISE,一种基于想象的可扩展机器人强化学习框架。核心是一个可组合的世界模型,能够(i)通过可控动力学模型预测多视角未来状态,(ii)利用进展值模型评估想象结果,生成有助于策略改进的信息性优势。这种可组合设计允许状态和价值分别采用最合适的独立架构与目标。这些组件集成进闭环自改进流程,持续生成想象轨迹、估计优势,并在想象空间中更新策略,无需昂贵的真实交互。在三个挑战性真实任务中,RISE相比已有方法取得显著提升:动态积木排序性能提高超过35%,背包打包提升45%,盒子闭合提升35%。

原文摘要 · Abstract (English)

Despite the sustained scaling on model capacity and data acquisition, Vision-Language-Action (VLA) models remain brittle in contact-rich and dynamic manipulation tasks, where minor execution deviations can compound into failures. While reinforcement learning (RL) offers a principled path to robustness, on-policy RL in the physical world is constrained by safety risk, hardware cost, and environment reset. To bridge this gap, we present RISE, a scalable framework of robotic reinforcement learning via imagination. At its core is a Compositional World Model that (i) predicts multi-view future via a controllable dynamics model, and (ii) evaluates imagined outcomes with a progress value model, producing informative advantages for the policy improvement. Such compositional design allows state and value to be tailored by best-suited yet distinct architectures and objectives. These components are integrated into a closed-loop self-improving pipeline that continuously generates imaginary rollouts, estimates advantages, and updates the policy in imaginary space without costly physical interaction. Across three challenging real-world tasks, RISE yields significant improvement over prior art, with more than +35% absolute performance increase in dynamic brick sorting, +45% for backpack packing, and +35% for box closing, respectively.

机器人学习强化学习想象推理自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。