用可控制世界模型实现高效强化微调,提升视觉语言动作模型鲁棒性。
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- 基于真实数据训练世界模型,模拟未来视觉观测并生成密集奖励
- 仅需400步微调即超越监督基线,样本效率高于传统仿真强化学习
- 适合需要高鲁棒性的机器人决策任务,尤其在分布外场景表现稳定
视觉-语言-动作(VLA)模型虽能实现具身决策,但过度依赖模仿学习导致误差累积且对分布外变化敏感。强化学习(RL)可缓解此问题,但通常需昂贵的真实交互或面临仿真到现实的差距。本文提出VLA-RFT,一种利用数据驱动世界模型作为可控模拟器的强化微调框架。该模型从真实交互数据中训练,能根据动作预测未来视觉观测,从而在策略回放中生成基于目标达成参考的密集轨迹级奖励。这一设计提供高效且动作对齐的学习信号,大幅降低样本需求。仅需不到400次微调步骤,VLA-RFT即超越强监督基线,并在效率上优于基于仿真强化学习的方法。此外,其在扰动条件下仍保持稳定任务执行能力。结果表明,基于世界模型的强化微调是提升VLA模型泛化性与鲁棒性的实用后训练范式。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models enable embodied decision-making but rely heavily on imitation learning, leading to compounding errors and poor robustness under distribution shift. Reinforcement learning (RL) can mitigate these issues yet typically demands costly real-world interactions or suffers from sim-to-real gaps. We introduce VLA-RFT, a reinforcement fine-tuning framework that leverages a data-driven world model as a controllable simulator. Trained from real interaction data, the simulator predicts future visual observations conditioned on actions, allowing policy rollouts with dense, trajectory-level rewards derived from goal-achieving references. This design delivers an efficient and action-aligned learning signal, drastically lowering sample requirements. With fewer than 400 fine-tuning steps, VLA-RFT surpasses strong supervised baselines and achieves greater efficiency than simulator-based RL. Moreover, it exhibits strong robustness under perturbed conditions, sustaining stable task execution. Our results establish world-model-based RFT as a practical post-training paradigm to enhance the generalization and robustness of VLA models. For more details, please refer to https://vla-rft.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。