通过多模态传感器提升复杂操作任务的演示数据质量与分割精度
TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks
- 在机械手集成视觉、力矩、位姿等多模态传感器,实现同步采集
- 在电缆安装任务中达到90%以上分割准确率,多模态提升显著
- 适合需要精细物理交互的机器人操作任务研究者使用
任务分解对理解与学习复杂长时程操作任务至关重要。尤其在涉及丰富物理交互的任务中,仅依赖视觉观测和机器人本体感知常无法揭示底层事件转换。这要求高效收集高质量多模态数据,并具备鲁棒的分割方法以将示范分解为有意义的模块。基于手持式通用操作接口(UMI)理念,我们提出TacUMI,一种集成视觉触觉传感器、力矩传感器和姿态追踪器的紧凑型、兼容机器人的夹持器设计,可在人类示范过程中同步获取所有模态数据。随后,我们提出一种多模态分割框架,利用时间模型检测序列操作中的语义事件边界。在具有挑战性的电缆安装任务上的评估显示,分割准确率超过90%,且更多模态带来显著性能提升,验证了TacUMI在接触丰富任务中可扩展的多模态示范数据采集与分割方面的实用基础。
原文摘要 · Abstract (English)
Task decomposition is critical for understanding and learning complex long-horizon manipulation tasks. Especially for tasks involving rich physical interactions, relying solely on visual observations and robot proprioceptive information often fails to reveal the underlying event transitions. This raises the requirement for efficient collection of high-quality multi-modal data as well as robust segmentation method to decompose demonstrations into meaningful modules. Building on the idea of the handheld demonstration device Universal Manipulation Interface (UMI), we introduce TacUMI, a multi-modal data collection system that integrates additionally ViTac sensors, force-torque sensor, and pose tracker into a compact, robot-compatible gripper design, which enables synchronized acquisition of all these modalities during human demonstrations. We then propose a multi-modal segmentation framework that leverages temporal models to detect semantically meaningful event boundaries in sequential manipulations. Evaluation on a challenging cable mounting task shows more than 90 percent segmentation accuracy and highlights a remarkable improvement with more modalities, which validates that TacUMI establishes a practical foundation for both scalable collection and segmentation of multi-modal demonstrations in contact-rich tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。