arXiv:2511.17001cs.RO2025-11被引 2

统一机器人动作的相机坐标系表示,实现跨平台动作语义一致

Unify Robot Actions in Camera Frame

  • 用相机外参将不同机器人的动作转为统一相机帧下的标准表示
  • 在16个数据集上生成约9.7万条校准后的动作数据,提升跨平台训练效果
  • 无需训练、不依赖具体机器人,适用于真实与仿真环境

跨体感机器人学习需要在不同机器人平台上保持动作表示的语义一致性。现有方法或存在平台特异性差异,或依赖特定机器人头或隐空间学习,无法根本解决不匹配问题。本文提出在相机帧中统一机器人动作,利用相机外参使动作具备一致的几何语义,适用于单臂与双臂机器人。然而多数数据集缺乏相机外参标注,现有离线标定方法易陷入局部极小值或需机器人特定训练数据。为此,我们提出 CalibAll:一种无需训练、机器人无关的标注流程,可对离线数据集估计相机外参,并将异构机器人动作转换为标准化的相机帧动作。CalibAll 采用粗到精标定策略:先通过时间性PnP获得稳定初始化,再通过可微渲染进行高精度优化。除外参外,CalibAll 还生成标准化的TCP姿态动作及辅助标注。我们在4种机器人平台的16个数据集上应用该方法,生成约97,000条校准数据。下游仿真与真实机器人实验表明,使用相机帧动作进行跨体感预训练可达到当前最优性能。

原文摘要 · Abstract (English)

Cross-embodiment robot learning requires a unified action representation with consistent semantics across robot platforms. Existing representations suffer from platform-specific inconsistencies, while current solutions either maintain embodiment-specific action heads or learn latent action spaces, without fundamentally resolving the mismatch. We propose to unify robot actions in the camera frame using camera extrinsics, so that actions share consistent geometric semantics across different robot embodiments, including both single-arm and bimanual robots. However, most existing datasets lack camera extrinsic annotations, and existing offline calibration methods either suffer from local minima or require robot-specific training data. To address this gap, we present CalibAll, a training-free, robot-independent annotation pipeline that estimates camera extrinsics for offline datasets and converts heterogeneous robot actions into standardized camera-frame actions. CalibAll follows a coarse-to-fine calibration strategy: temporal PnP provides a stable initialization, followed by differentiable rendering-based refinement for high precision. Beyond extrinsics, CalibAll produces standardized TCP-pose actions and auxiliary annotations. We apply CalibAll to 16 datasets across 4 robot platforms, producing approximately 97K calibrated data episodes. Downstream simulation and real-robot experiments show that cross-embodiment pretraining with camera-frame actions achieves state-of-the-art performance.

机器人学习动作统一相机外参数据标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。