arXiv:2511.00091cs.CVcs.RO2025-11被引 66

用残差强化学习自动生成数据,让视觉语言动作模型越用越强。

Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

  • 通过轻量级残差智能体探测主模型失败区域,定位改进点。
  • 在仿真和真实机械臂上实现99%任务成功率,部分场景提升超50%。
  • 无需人工标注,适合希望自动化优化机器人模型的团队使用。

监督微调(SFT)已成为大型视觉-语言-动作(VLA)模型的标准后训练策略,但其依赖昂贵的人工示范,限制了可扩展性和泛化能力。我们提出探查、学习、蒸馏(PLD)三阶段即插即用框架,通过残差强化学习与分布感知的数据采集来改进VLA模型。第一阶段训练轻量级残差执行器,探测通用模型的失败区域;第二阶段采用混合回放机制,在对齐通用模型部署分布的同时捕捉恢复行为;第三阶段将筛选后的轨迹通过标准SFT回填至通用模型。PLD在LIBERO上达到接近饱和的99%任务成功率,在SimplerEnv上提升超50%,并在真实世界Franka与YAM机械臂操作任务中实现100%成功。消融实验表明,残差探测与分布感知回放是获取部署对齐数据的关键,显著提升已知与未知任务表现,为构建自进化VLA模型提供可扩展路径。

原文摘要 · Abstract (English)

Supervised fine-tuning (SFT) has become the de facto post-training strategy for large vision-language-action (VLA) models, but its reliance on costly human demonstrations limits scalability and generalization. We propose Probe, Learn, Distill (PLD), a three-stage plug-and-play framework that improves VLAs through residual reinforcement learning (RL) and distribution-aware data collection. In Stage 1, we train lightweight residual actors to probe failure regions of the VLA generalist. In Stage 2, we use a hybrid rollout scheme that aligns collected trajectories with the generalist's deployment distribution while capturing recovery behaviors. In Stage 3, we distill the curated trajectories back into the generalist with standard SFT. PLD achieves near-saturated 99% task success on LIBERO, over 50% gains in SimplerEnv, and 100% success on real-world Franka and YAM arm manipulation tasks. Ablations show that residual probing and distribution-aware replay are key to collecting deployment-aligned data that improves both seen and unseen tasks, offering a scalable path toward self-improving VLA models.

视觉语言动作强化学习机器人自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。