Z-1用强化学习提升视觉语言动作模型,不依赖私有数据也能显著提高机器人操作成功率。
Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models

- 基于流模型的VLA,采用分任务组相对策略优化提升性能
- 24个RoboCasa任务平均成功率达80.6%,比初始SFT提升13.2个百分点
- 仅用公开数据即可实现超越现有最优模型,适合追求高效部署的机器人研究者
视觉语言动作(VLA)模型通过连接语言指令、视觉观测与连续控制,为机器人操作提供了有前景的框架。然而,现有策略大多受限于行为克隆或从固定演示中进行监督微调(SFT),难以从自身失败中学习改进。本文提出Z-1,一种针对流模型VLA的强化学习后训练框架。Z-1基于$π_{0.5}$,仅使用公开的RoboCasa演示进行SFT,随后在24个标准RoboCasa任务上采用任务级组相对策略优化(GRPO)。为提升在线优化效率与稳定性,Z-1结合共享前缀回溯构造、树状轨迹分支、完成感知奖励校准及选择性联合训练视觉语言模型(VLM)与动作专家。在全部24个任务上,Z-1平均成功率达到80.6%,相比SFT初始化提升13.2个百分点,并优于已发表的最先进模型。结果表明,系统性的GRPO后训练可显著提升流模型VLA策略,且无需额外私有演示。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observations, and continuous control. However, most existing policies remain limited by behavior cloning or supervised fine-tuning (SFT) from fixed demonstrations, which provides limited opportunity to improve from the policy's own failures. In this paper, we present Z-1, a reinforcement learning (RL) post-training framework for flow-based VLA models. Built on top of $π_{0.5}$, Z-1 uses only publicly released RoboCasa demonstrations for SFT and then applies a task-wise Group Relative Policy Optimization (GRPO) strategy across $24$ standard RoboCasa tasks. To improve the efficiency and stability of online optimization, Z-1 combines shared-prefix rollout construction, tree-structured trajectory branching, completion-aware reward calibration, and selective joint training of VLM and Action Expert. Across all $24$ RoboCasa tasks, Z-1 achieves an average success rate of $80.6\%$, improving over its SFT initialization by $13.2\%$ points and outperforms the published sota models. These results show that systematic GRPO post-training can substantially improve flow-based VLA policies without additional private demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。