arXiv:2505.19789cs.LG2025-05NeurIPS被引 104

用强化学习提升视觉语言动作模型的泛化能力

What Can RL Bring to VLA Generalization? An Empirical Study

论文配图:What Can RL Bring to VLA Generalization? An Empirical Study
图 1 · 摘自论文原文
  • 对比监督微调,强化学习通过试错优化任务目标
  • PPO算法显著提升语义理解与执行鲁棒性,视觉表现相当
  • 提出高效PPO训练方案,适合实际部署的VLA改进

大型视觉语言动作(VLA)模型在具身智能中展现出巨大潜力。然而,其主流的监督微调(SFT)方法在分布外场景下易受累积误差影响,限制了泛化能力。强化学习(RL)通过试错优化任务目标,为突破此瓶颈提供了可能,但其对VLA泛化效果的系统性理解仍不足。为此,我们构建了一个全面的VLA泛化评估基准,系统研究了不同视觉、语义和执行维度下强化学习微调的影响。大量实验表明,尤其采用PPO算法时,强化学习微调显著提升了语义理解与执行鲁棒性,同时保持与SFT相当的视觉鲁棒性。我们发现PPO比基于大语言模型的DPO、GRPO等方法更适用于VLA。此外,我们提出了一个高效的PPO训练实用方案,并验证了其在提升VLA泛化中的实际价值。

原文摘要 · Abstract (English)

Large Vision-Language Action (VLA) models have shown significant potential for embodied AI. However, their predominant training via supervised fine-tuning (SFT) limits generalization due to susceptibility to compounding errors under distribution shifts. Reinforcement learning (RL) offers a path to overcome these limitations by optimizing for task objectives via trial-and-error, yet a systematic understanding of its specific generalization benefits for VLAs compared to SFT is lacking. To address this, our study introduces a comprehensive benchmark for evaluating VLA generalization and systematically investigates the impact of RL fine-tuning across diverse visual, semantic, and execution dimensions. Our extensive experiments reveal that RL fine-tuning, particularly with PPO, significantly enhances generalization in semantic understanding and execution robustness over SFT, while maintaining comparable visual robustness. We identify PPO as a more effective RL algorithm for VLAs than LLM-derived methods like DPO and GRPO. We also develop a simple recipe for efficient PPO training on VLAs, and demonstrate its practical utility for improving VLA generalization. The project page is at https://rlvla.github.io

强化学习视觉语言动作泛化能力PPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。