arXiv:2511.20633cs.RO2025-11被引 19

用预训练世界模型和稳定强化学习,让机器人动作策略更鲁棒高效。

Reinforcing Action Policies by Prophesying

  • 基于动作-视频的统一世界模型,实现跨场景快速适配。
  • 引入梯度重加权机制,使强化学习训练更稳定,提升成功率5-17%。
  • 适合需要低成本调优真实机器人策略的研究者与开发者。

视觉-语言-动作(VLA)策略在对齐语言、感知与机器人控制方面表现优异,但多数VLA仅通过模仿学习训练,易过拟合示范数据,在分布偏移下表现脆弱。强化学习(RL)可直接优化任务奖励,但真实机器人交互成本高,传统仿真器难构建且难以迁移。本文提出通过预训练的世界模型与适配流式动作头的强化学习方法,解决后训练阶段的数据效率与优化稳定性问题。首先引入Prophet——一个在大规模异构机器人数据上预训练的动作-视频世界模型,可快速少样本适配新机器人、物体与环境,生成可用的仿真滚动数据。在此基础上,提出FlowScale,将流式梯度策略优化(Flow-GRPO)与逐步内在重加权结合,稳定梯度更新。实验表明,该方案在公开基准上提升5-17%成功率,在真实机器人上提升24-30%,适用于多种VLA骨干网络。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) policies excel in aligning language, perception, and robot control. However, most VLAs are trained purely by imitation, which overfits to demonstrations, and is brittle under distribution shift. Reinforcement learning (RL) directly optimizes task reward and thus addresses this misalignment, but real-robot interaction is expensive and conventional simulators are hard to engineer and transfer. We address both data efficiency and optimization stability in VLA post-training via a learned world model and an RL procedure tailored to flow-based action heads. Specifically, we first introduce Prophet, a unified action-to-video robot world model pretrained on large-scale, heterogeneous robot data to learn reusable action-outcome dynamics and then few-shot adapted to new robots, objects, and environments, yielding a rollout-ready simulator. Upon Prophet, we reinforce action policies with our proposed FlowScale, which couples Flow-GRPO with intrinsic stepwise reweighting to stabilize gradients. Together, our solution provides a practical, data- and compute-efficient path to VLA post-training. Experiments show 5-17% success gains on public benchmarks and 24-30% on real robots across diverse VLA backbones.

机器人强化学习世界模型动作策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。