arXiv:2604.14089cs.ROcs.AI2026-04被引 4

给可穿戴机械臂数据采集系统加了激光雷达,让机器人在复杂环境里更稳地干活。

UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

论文配图:UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception
图 1 · 摘自论文原文
  • 在腕戴设备上集成低成本激光雷达,实现3D空间感知下的稳定定位
  • 实测在变形物体和活动场景中成功率显著提升,能完成原版做不到的任务
  • 开源全流程工具,适合做具身智能数据采集的研究者使用

我们提出UMI-3D,是通用操作接口(UMI)的多模态扩展,用于在具身操作中实现鲁棒且可扩展的数据采集。尽管UMI支持便携式腕戴数据采集,但其依赖单目视觉SLAM,在遮挡、动态场景和跟踪失败时表现脆弱,限制了在真实环境中的应用。UMI-3D通过将轻量级低成本的激光雷达紧密集成到腕戴界面中,实现了以激光雷达为中心的SLAM,可在复杂条件下实现精确的度量尺度位姿估计。我们进一步开发了硬件同步的多模态感知流水线和统一时空标定框架,对齐视觉观测与激光点云,生成一致的演示三维表示。尽管保持原有的二维视觉运动策略形式,UMI-3D显著提升了采集数据的质量与可靠性,直接带来策略性能提升。大量真实世界实验表明,UMI-3D不仅在标准操作任务中取得高成功率,还实现了原版仅靠视觉的UMI难以或无法完成的任务,包括大形变物体操作和可动部件操控。该系统支持从数据采集、对齐、训练到部署的端到端流程,同时保留了原始UMI的便携性与易用性。所有软硬件组件均已开源,以促进大规模数据采集并加速具身智能研究:https://umi-3d.github.io。

原文摘要 · Abstract (English)

We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation. While UMI enables portable, wrist-mounted data acquisition, its reliance on monocular visual SLAM makes it vulnerable to occlusions, dynamic scenes, and tracking failures, limiting its applicability in real-world environments. UMI-3D addresses these limitations by introducing a lightweight and low-cost LiDAR sensor tightly integrated into the wrist-mounted interface, enabling LiDAR-centric SLAM with accurate metric-scale pose estimation under challenging conditions. We further develop a hardware-synchronized multimodal sensing pipeline and a unified spatiotemporal calibration framework that aligns visual observations with LiDAR point clouds, producing consistent 3D representations of demonstrations. Despite maintaining the original 2D visuomotor policy formulation, UMI-3D significantly improves the quality and reliability of collected data, which directly translates into enhanced policy performance. Extensive real-world experiments demonstrate that UMI-3D not only achieves high success rates on standard manipulation tasks, but also enables learning of tasks that are challenging or infeasible for the original vision-only UMI setup, including large deformable object manipulation and articulated object operation. The system supports an end-to-end pipeline for data acquisition, alignment, training, and deployment, while preserving the portability and accessibility of the original UMI. All hardware and software components are open-sourced to facilitate large-scale data collection and accelerate research in embodied intelligence: \href{https://umi-3d.github.io}{https://umi-3d.github.io}.

具身智能数据采集激光雷达机械臂

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。