让视觉语言动作模型更稳定可靠,通过多样化经验训练提升真实机器人任务表现
Scaling by Diversified Experience for Vision-Language-Action Models

- 分离意图与控制特征,避免高阶推理干扰低阶操作
- 采用相似样本引导强化学习,显著减少策略更新波动
- 在真实机器人任务中成功率更高,适合部署于复杂场景
视觉-语言-动作(VLA)模型在实际部署中面临高阶推理与低阶控制耦合、策略优化不稳定的挑战。本文提出SyVLA,一种基于多样化经验训练的鲁棒VLA模型。通过引入意图解耦算法,将控制相关特征从推理上下文中分离;设计相似样本引导的强化学习流程,稳定策略更新并缓解分布偏移问题。在真实机器人任务及多模态基准上的大量实验表明,SyVLA相比现有方法实现了更高的任务成功率和更强的分布外泛化能力,同时有效保留了核心视觉-语言能力。代码与数据集已公开于项目主页。
原文摘要 · Abstract (English)
Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. In this paper, we introduce SyVLA, a robust VLA model trained with diversified experiences. We propose an Intention Decoupling algorithm to isolate control-relevant features from reasoning contexts and a similar-sample guided RL pipeline to stabilize policy updates and mitigate distribution shift. Extensive experiments on real-world robotic tasks and multi-modal benchmarks demonstrate that SyVLA achieves superior task success rates and stronger out-of-distribution generalization compared to existing methods, while effectively preserving core vision-language capabilities. Codes and Datasets is released on \href{https://sy-vla.github.io/}{project page}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。