arXiv:2606.19784cs.RO2026-06

让机器人视觉-语言-动作模型自动适应旋转,提升泛化能力。

EquiVLA: A General Framework for Rotationally Equivariant Vision-Language-Action Models

论文配图:EquiVLA: A General Framework for Rotationally Equivariant Vision-Language-Action Models
图 1 · 摘自论文原文
  • 通过等变感知与动作头构建旋转等变链,实现端到端旋转不变性。
  • 在LIBERO上平均成功率92.6%(基线78.1%),真实机器人任务成功率达72%。
  • 适用于任意视觉语言主干+扩散动作头结构,适合需要旋转鲁棒性的机器人研究。

视觉-语言-动作(VLA)模型已成为通用机器人操作的强大范式,但缺乏几何归纳偏置:在特定朝向训练的策略需大量数据才能跨旋转配置泛化。本文提出 extsc{EquiVLA},首个面向端到端 $ m SO(2)$ 等变 VLA 模型的通用框架,适用于任意冻结视觉语言主干与流匹配扩散变换器动作头的组合。 extsc{EquiVLA} 引入 extsc{EquiPerceptor},从冻结的 ViT 特征中生成近似 $ m SO(2)$ 等变视觉表征;以及 extsc{EquiActor},一个精确 $ m SO(2)$ 等变的流匹配扩散变换器动作头。二者共同构建从相机观测到预测动作序列的近似 $ m SO(2)$ 等变链。在 GR00T~N1.5 上实例化,并在四个 LIBERO 套件、CALVIN ABCD→D 以及 Mobile ALOHA 上的五个真实机器人任务上评估, extsc{EquiVLA} 在 LIBERO 上实现 92.6% 的平均成功率(基线 78.1%),在 CALVIN 上平均序列长度达 4.03(基线 3.45),真实机器人任务成功率从 54% 提升至 72%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for generalist robot manipulation, yet they lack geometric inductive biases: policies trained at specific orientations require substantially more data to generalize across rotational configurations. We present \textsc{EquiVLA}, the first general framework for end-to-end $\mathrm{SO}(2)$-equivariant VLA models, applicable to any architecture coupling a frozen vision-language backbone with a flow-matching Diffusion Transformer action head. \textsc{EquiVLA} introduces \textsc{EquiPerceptor}, which produces approximately $\mathrm{SO}(2)$-equivariant visual representations from frozen ViT features; and \textsc{EquiActor}, an exactly $\mathrm{SO}(2)$-equivariant flow-matching Diffusion Transformer action head. Together, they establish an approximate $\mathrm{SO}(2)$ equivariance chain from camera observations to predicted action sequences. Instantiated on GR00T~N1.5 and evaluated across four LIBERO suites, CALVIN ABCD$\to$D, and five real-robot tasks on Mobile ALOHA, \textsc{EquiVLA} achieves $92.6\%$ average success on LIBERO (vs. $78.1\%$ baseline), an average sequence length of $4.03$ on CALVIN (vs. $3.45$), and improves real-robot success from $54\%$ to $72\%$.

机器人等变网络视觉语言动作扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。