arXiv:2602.19710cs.CVcs.LG2026-02中稿 · Robotics: Science …被引 8

用统一姿态表征提升机器人视觉-语言-动作模型的泛化能力

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

  • 分离预训练与对齐阶段,用离散姿态令牌建模通用3D空间先验
  • 在RoboTwin 2.0上达79.5%成功率,LIBERO上达96.0%
  • 仅需每任务100次示范即可实现实体机器人跨物体泛化

现有视觉-语言-动作(VLA)模型因将高层感知与稀疏的具身动作监督耦合,常出现特征坍缩和训练效率低的问题。由于这些模型依赖面向视觉问答(VQA)优化的VLM骨干网络,虽擅长语义识别,却忽略决定不同动作模式的细微3D状态变化。为此,我们提出Pose-VLA,一种解耦范式:将VLA训练分为两阶段——首先在统一相机中心空间中通过姿态预训练提取通用3D空间先验,再在具体机器人动作空间中高效进行具身对齐。通过引入离散姿态令牌作为通用表示,Pose-VLA可无缝融合多源3D数据的空间定位信息与机器人示范的几何级轨迹。该框架采用两阶段预训练流程,先建立基于姿态的空间基础,再通过轨迹监督实现运动对齐。大量实验表明,Pose-VLA在RoboTwin 2.0上达到79.5%平均成功率,在LIBERO上表现优异达96.0%。真实世界实验进一步验证了其仅用每任务100次示范即可在多样物体间实现鲁棒泛化,证明了预训练范式的高效性。

原文摘要 · Abstract (English)

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA), they excel at semantic identification but often overlook subtle 3D state variations that dictate distinct action patterns. To resolve these misalignments, we propose Pose-VLA, a decoupled paradigm that separates VLA training into a pre-training phase for extracting universal 3D spatial priors in a unified camera-centric space, and a post-training phase for efficient embodiment alignment within robot-specific action space. By introducing discrete pose tokens as a universal representation, Pose-VLA seamlessly integrates spatial grounding from diverse 3D datasets with geometry-level trajectories from robotic demonstrations. Our framework follows a two-stage pre-training pipeline, establishing fundamental spatial grounding via poses followed by motion alignment through trajectory supervision. Extensive evaluations demonstrate that Pose-VLA achieves state-of-the-art results on RoboTwin 2.0 with a 79.5% average success rate and competitive performance on LIBERO at 96.0%. Real-world experiments further showcase robust generalization across diverse objects using only 100 demonstrations per task, validating the efficiency of our pre-training paradigm.

机器人学习姿态表征多模态预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。