arXiv:2604.10647cs.RO2026-04被引 5

用多模态感知让机器人学会真实触觉操作,像人一样摸着干。

OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction

论文配图:OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction
图 1 · 摘自论文原文
  • 手持设备同步采集视觉、触觉、力矩等多模态信号
  • 在抓取、擦拭等任务中实现高精度力控与触觉反馈
  • 适合需要精细接触控制的机器人操作研究者

UMI式交互接口虽能实现可扩展的机器人学习,但现有系统仍以视觉-运动为主,仅依赖RGB图像和轨迹,对物理交互信号获取有限。这在接触密集型操作中成为根本瓶颈,因成功依赖于难以仅从视觉推断的接触动力学,如触觉交互、内部抓握力和外部作用力矩。我们提出OmniUMI,一种通过人机对齐多模态交互实现物理基础机器人学习的统一框架。OmniUMI在紧凑手持设备中同步采集RGB、深度、轨迹、触觉传感、内部抓握力和外部作用力矩,并通过共享实体设计保持数据采集与部署一致性。为支持人机对齐示范,该系统通过双臂夹持器反馈,实现对内部抓握力、外部作用力矩和触觉交互的自然感知与调节。基于此接口,我们扩展了融合视觉、触觉及力相关观测的扩散策略,并通过阻抗执行实现运动与接触行为的统一调控。实验表明,在力敏感抓放、交互式表面擦除和触觉引导选择性释放任务中均表现可靠。总体而言,OmniUMI结合物理基础多模态数据采集与人机对齐交互,为接触密集型操作学习提供了可扩展基础。

原文摘要 · Abstract (English)

UMI-style interfaces enable scalable robot learning, but existing systems remain largely visuomotor, relying primarily on RGB observations and trajectory while providing only limited access to physical interaction signals. This becomes a fundamental limitation in contact-rich manipulation, where success depends on contact dynamics such as tactile interaction, internal grasping force, and external interaction wrench that are difficult to infer from vision alone. We present OmniUMI, a unified framework for physically grounded robot learning via human-aligned multimodal interaction. OmniUMI synchronously captures RGB, depth, trajectory, tactile sensing, internal grasping force, and external interaction wrench within a compact handheld system, while maintaining collection--deployment consistency through a shared embodiment design. To support human-aligned demonstration, OmniUMI enables natural perception and modulation of internal grasping force, external interaction wrench, and tactile interaction through bilateral gripper feedback and the handheld embodiment. Built on this interface, we extend diffusion policy with visual, tactile, and force-related observations, and deploy the learned policy through impedance-based execution for unified regulation of motion and contact behavior. Experiments demonstrate reliable sensing and strong downstream performance on force-sensitive pick-and-place, interactive surface erasing, and tactile-informed selective release. Overall, OmniUMI combines physically grounded multimodal data acquisition with human-aligned interaction, providing a scalable foundation for learning contact-rich manipulation.

机器人操作多模态感知力控人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。