arXiv:2602.12099cs.CV2026-02被引 7

用世界模型强化学习训练机器人,让其更聪明地完成复杂操作。

GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning

  • 基于世界模型的强化学习提升动作预测能力
  • 在折叠衣物等任务上性能提升约30%
  • 适合研究机器人智能与长时序决策的学者

视觉-语言-动作(VLA)模型直接从当前观测预测多步动作序列,受限于场景理解不足和未来预测能力弱。相比之下,预训练于网络规模视频数据的视频世界模型具备强大的时空推理和精准未来预测能力,是增强VLA学习的理想基础。为此,我们提出GigaBrain-0.5M*,一种通过世界模型引导的强化学习训练的VLA模型。该模型建立在GigaBrain-0.5之上,后者在超过10,000小时机器人操作数据上预训练,其中间版本目前在国际RoboChallenge基准上排名第一。GigaBrain-0.5M*进一步通过RAMP(基于世界模型条件策略的强化学习)集成世界模型引导的强化学习,实现稳健的跨任务适应。实验证明,RAMP相比RECAP基线在挑战性任务如 exttt{Laundry Folding}、 exttt{Box Packing}和 exttt{Espresso Preparation}上取得约30%的性能提升。关键的是,GigaBrain-0.5M*展现出可靠的长时序执行能力,在真实部署视频中持续完成复杂操作且未出现失败。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models that directly predict multi-step action chunks from current observations face inherent limitations due to constrained scene understanding and weak future anticipation capabilities. In contrast, video world models pre-trained on web-scale video corpora exhibit robust spatiotemporal reasoning and accurate future prediction, making them a natural foundation for enhancing VLA learning. Therefore, we propose \textit{GigaBrain-0.5M*}, a VLA model trained via world model-based reinforcement learning. Built upon \textit{GigaBrain-0.5}, which is pre-trained on over 10,000 hours of robotic manipulation data, whose intermediate version currently ranks first on the international RoboChallenge benchmark. \textit{GigaBrain-0.5M*} further integrates world model-based reinforcement learning via \textit{RAMP} (Reinforcement leArning via world Model-conditioned Policy) to enable robust cross-task adaptation. Empirical results demonstrate that \textit{RAMP} achieves substantial performance gains over the RECAP baseline, yielding improvements of approximately 30\% on challenging tasks including \texttt{Laundry Folding}, \texttt{Box Packing}, and \texttt{Espresso Preparation}. Critically, \textit{GigaBrain-0.5M$^*$} exhibits reliable long-horizon execution, consistently accomplishing complex manipulation tasks without failure as validated by real-world deployment videos on our \href{https://gigabrain05m.github.io}{project page}.

机器人世界模型强化学习长时序决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。