提升机器人视觉-语言-动作模型的空间感知与强化学习稳定性
OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL

- 采用混合变换器架构融合推理、空间与动作专家
- 在LIBERO基准上超越主流方法,提升动作精度与训练稳定性
- 适合研究具身智能与机器人控制的开发者参考
视觉-语言-动作(VLA)模型代表了具身智能的新范式,但现有框架常面临空间感知不精确、多模态融合不佳及强化学习不稳定等问题。为此,我们提出OmniVLA-RL,一种基于混合变换器(MoT)设计的新架构,协同整合推理、空间与动作专家。此外,我们引入Flow-GSPO,将流匹配重新表述为随机微分方程(SDE)过程,并结合分组分割策略优化(GSPO),以提升动作精度与训练鲁棒性。在LIBERO与LIBERO-Plus基准上的大量评估表明,OmniVLA-RL实现了良好的整体性能,超越主流现有方法,有效克服了当前VLA模型的根本局限。
原文摘要 · Abstract (English)
Visual-Language-Action (VLA) models represent a paradigm shift in embodied AI, yet existing frameworks often struggle with imprecise spatial perception, suboptimal multimodal fusion, and instability in reinforcement learning. To bridge these gaps, we propose OmniVLA-RL, a novel architecture that leverages a Mix-of-Transformers (MoT) design to synergistically integrate reasoning, spatial, and action experts. Furthermore, we introduce Flow-GSPO, which reformulates flow matching as a Stochastic Differential Equation (SDE) process and integrates it with Group Segmented Policy Optimization (GSPO) to enhance action precision and training robustness. Extensive evaluations on the LIBERO and LIBERO-Plus benchmarks demonstrate that OmniVLA-RL achieves decent overall performance and surpasses mainstream existing methods, effectively overcoming the fundamental limitations of current VLA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。