arXiv:2604.17706cs.RO2026-04

提升机器人视觉-语言-动作模型的空间感知与强化学习稳定性

OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL

论文配图:OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL
图 1 · 摘自论文原文
  • 采用混合变换器架构融合推理、空间与动作专家
  • 在LIBERO基准上超越主流方法,提升动作精度与训练稳定性
  • 适合研究具身智能与机器人控制的开发者参考

视觉-语言-动作(VLA)模型代表了具身智能的新范式,但现有框架常面临空间感知不精确、多模态融合不佳及强化学习不稳定等问题。为此,我们提出OmniVLA-RL,一种基于混合变换器(MoT)设计的新架构,协同整合推理、空间与动作专家。此外,我们引入Flow-GSPO,将流匹配重新表述为随机微分方程(SDE)过程,并结合分组分割策略优化(GSPO),以提升动作精度与训练鲁棒性。在LIBERO与LIBERO-Plus基准上的大量评估表明,OmniVLA-RL实现了良好的整体性能,超越主流现有方法,有效克服了当前VLA模型的根本局限。

原文摘要 · Abstract (English)

Visual-Language-Action (VLA) models represent a paradigm shift in embodied AI, yet existing frameworks often struggle with imprecise spatial perception, suboptimal multimodal fusion, and instability in reinforcement learning. To bridge these gaps, we propose OmniVLA-RL, a novel architecture that leverages a Mix-of-Transformers (MoT) design to synergistically integrate reasoning, spatial, and action experts. Furthermore, we introduce Flow-GSPO, which reformulates flow matching as a Stochastic Differential Equation (SDE) process and integrates it with Group Segmented Policy Optimization (GSPO) to enhance action precision and training robustness. Extensive evaluations on the LIBERO and LIBERO-Plus benchmarks demonstrate that OmniVLA-RL achieves decent overall performance and surpasses mainstream existing methods, effectively overcoming the fundamental limitations of current VLA models.

具身智能视觉-语言-动作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。