arXiv:2510.01607cs.ROcs.CV2025-10被引 25

用头戴设备记录人类操作时的视线动作,让机器人学会主动观察来完成复杂双手操作。

ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations

  • 通过头戴显示设备捕捉操作者头部运动,建立视觉注意力与操作之间的联系。
  • 在6个任务中平均成功率70%,新物体和新环境仍保持56%成功率。
  • 便携式系统可高效收集高质量数据,适合想快速部署机器人的团队。

我们提出ActiveUMI,一个用于将真实世界中人类示范迁移至具备复杂双臂操作能力机器人的数据采集框架。ActiveUMI结合便携式VR遥操作系统与带传感器的控制器,精确对齐人机运动学。为保证移动性与数据质量,引入沉浸式3D建模、自包含可穿戴计算机及高效校准方法。其核心在于捕获主动的、第一人称视角的感知信息:通过头戴显示器记录操作者的主动头部运动,学习视觉注意力与操作之间的关键关联。我们在六个具有挑战性的双臂任务上评估了ActiveUMI。仅使用ActiveUMI数据训练的策略,在分布内任务上达到平均70%的成功率,并展现出强泛化能力,在新物体和新环境中仍保持56%的成功率。结果表明,结合学习到的主动感知,便携式数据采集系统为构建通用且高性能的真实世界机器人策略提供了一条有效且可扩展的路径。

原文摘要 · Abstract (English)

We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples a portable VR teleoperation kit with sensorized controllers that mirror the robot's end-effectors, bridging human-robot kinematics via precise pose alignment. To ensure mobility and data quality, we introduce several key techniques, including immersive 3D model rendering, a self-contained wearable computer, and efficient calibration methods. ActiveUMI's defining feature is its capture of active, egocentric perception. By recording an operator's deliberate head movements via a head-mounted display, our system learns the crucial link between visual attention and manipulation. We evaluate ActiveUMI on six challenging bimanual tasks. Policies trained exclusively on ActiveUMI data achieve an average success rate of 70\% on in-distribution tasks and demonstrate strong generalization, retaining a 56\% success rate when tested on novel objects and in new environments. Our results demonstrate that portable data collection systems, when coupled with learned active perception, provide an effective and scalable pathway toward creating generalizable and highly capable real-world robot policies.

机器人操控主动感知虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。