arXiv:2608.26058cs.RO2026-08

统一异构机器人动作空间,让不同形态机器人共享同一套操作策略。

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

论文配图:One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation
图 1 · 摘自论文原文
  • 以摄像头视角定义通用动作几何,将不同机器人动作对齐到统一空间。
  • 单个模型在多个数据集上达到98.3%~62.0%准确率,无需微调。
  • 适合多机器人协同、跨平台通用策略开发场景。

规模化通用视觉-语言-动作(VLA)策略面临的核心瓶颈是具身数据的固有异质性,涵盖多样化的机器人形态、相机配置和底层动作空间。现有方法通常依赖显式的动作重映射、人机视频合成或针对特定数据集的适配分支,从根本上阻碍了统一策略的联合学习。本文提出UCAG-P,一种以摄像头为中心的统一动作表述方法,通过结构化对齐将异构具身数据纳入共享几何动作空间。不同于将机器人特异性指令作为统一策略目标,UCAG-P通过图像与相机坐标系中的可观察锚点运动来表示操作,将机械臂、人形机器人及人手视为同一动作范式下的不同体现。一个几何条件的动作翻译器结合预测运动与目标实体运动学,生成可执行控制信号。由此产生的解耦架构使共享VLA策略能够学习可迁移的操作几何,同时保持对具体实体的可控性。UCAG-P在4.03千小时机器人与仿真数据及2.34千小时人类示范数据上训练。单一检查点在LIBERO上达98.3%,RoboTwin Easy和Hard分别达88.7%和89.2%,LIBERO-Plus零样本达82.0%,RoboCasa GR-1达62.0%,均无需基准微调。

原文摘要 · Abstract (English)

Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.

具身智能动作对齐统一策略多机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。