arXiv:2605.13403cs.ROcs.CV2026-05被引 5

用连续旋转表示动作,让机器人模型更懂真实世界运动规律。

RotVLA: Rotational Latent Action for Vision-Language-Action Model

论文配图:RotVLA: Rotational Latent Action for Vision-Language-Action Model
图 1 · 摘自论文原文
  • 用SO(n)旋转空间建模连续动作,避免离散编码的缺陷。
  • 在LIBERO上达98.2%成功率,RoboTwin2.0上超88%且实机表现强。
  • 适合做具身智能、机器人控制的通用动作生成系统。

隐式动作模型(LAMs)已成为视觉-语言-动作(VLA)模型预训练中处理异构数据的有效范式,提供了跨不同机器人平台的统一动作空间。然而,现有LAMs多依赖离散量化编解码流程,易导致帧重建行为平凡、表征能力有限且缺乏物理意义结构。本文提出RotVLA,基于连续旋转隐式动作表示构建VLA框架。隐式动作被建模为SO(n)中的元素,具备连续性、可组合性及与真实动作动态一致的几何结构。引入三元组帧学习框架,强化有意义的时间动态并防止退化。RotVLA由视觉-语言模型主干和流匹配动作头组成,在大规模跨平台机器人数据集与人类视频上进行隐式动作监督预训练。下游机器人控制任务中,流匹配头扩展为统一动作专家,联合去噪隐式动作与机器人动作。其中,隐式动作作为潜在规划器,提供高层指导以条件化动作生成。仅使用1700+小时预训练数据与1.7B参数,RotVLA在LIBERO上达到98.2%成功率,在RoboTwin2.0的干净与随机设置下分别达89.6%与88.5%。在真实操控任务中也表现出色,持续优于现有VLA模型。

原文摘要 · Abstract (English)

Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretraining, offering a unified action space across embodiments. However, existing LAMs often rely on discrete quantization encode and decode pipelines, which can lead to trivial frame reconstruction behavior, limited representational capacity, and a lack of physically meaningful structure. We introduce RotVLA, a VLA framework built on a continuous rotational latent action representation. Latent actions are modeled as elements of SO(n), providing continuity, compositionality, and structured geometry aligned with real-world action dynamics. A triplet frame learning framework further enforces meaningful temporal dynamics while avoiding degeneration. RotVLA consists of a VLM backbone and a flow-matching action head, pretrained on large-scale cross-embodiment robotic datasets and human videos with latent-action supervision. For downstream robot control, the flow-matching head is extended into a unified action expert that jointly denoises latent and robot actions. Here, latent actions serve as a latent planner, providing high-level guidance that conditions action generation. With only 1.7B parameters and 1700+ hours of pretraining data, RotVLA achieves 98.2% on LIBERO and 89.6% / 88.5% on RoboTwin2.0 under clean and randomized settings, respectively. It also demonstrates strong real-world performance on manipulation tasks, consistently outperforming existing VLA models.

机器人控制隐式动作连续表示具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。