arXiv:2512.09851cs.ROcs.CV2025-12被引 13

新传感器+新算法让机器人同时感知触觉与视觉,抓取成功率提升至85.5%。

Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation

  • 开发可同步捕捉触觉与视觉信号的透明触觉传感器TacThru。
  • 在5个真实任务中实现85.5%平均成功率,优于纯触觉(66.3%)和纯视觉(55.4%)方案。
  • 适合需要精准接触判断与多模态协同的复杂抓取场景。

机器人操作需兼具丰富的多模态感知与有效的学习框架以应对复杂现实任务。透皮传感器(STS)融合触觉与视觉感知,具备潜力;现代模仿学习则为策略获取提供强大工具。然而,现有STS设计缺乏同步多模态感知,且触觉追踪不可靠。将这些丰富信号融入基于学习的操纵流程仍面临挑战。本文提出TacThru传感器,实现视觉与触觉信号的同时感知及稳健提取,并构建了基于Transformer的扩散策略的模仿学习框架TacThru-UMI。该传感器采用全透明弹性体、持续照明、新型关键线标记与高效追踪机制;学习系统通过融合多模态信号实现端到端策略生成。在五个具有挑战性的现实任务中,TacThru-UMI平均成功率达85.5%,显著高于仅触觉策略(66.3%)与仅视觉策略(55.4%)。系统在薄软物体接触检测及需多模态协同的精密操作中表现突出。本工作表明,结合同步多模态感知与现代学习框架,可实现更精确、自适应的机器人操作。

原文摘要 · Abstract (English)

Robotic manipulation requires both rich multimodal perception and effective learning frameworks to handle complex real-world tasks. See-through-skin (STS) sensors, which combine tactile and visual perception, offer promising sensing capabilities, while modern imitation learning provides powerful tools for policy acquisition. However, existing STS designs lack simultaneous multimodal perception and suffer from unreliable tactile tracking. Furthermore, integrating these rich multimodal signals into learning-based manipulation pipelines remains an open challenge. We introduce TacThru, an STS sensor enabling simultaneous visual perception and robust tactile signal extraction, and TacThru-UMI, an imitation learning framework that leverages these multimodal signals for manipulation. Our sensor features a fully transparent elastomer, persistent illumination, novel keyline markers, and efficient tracking, while our learning system integrates these signals through a Transformer-based Diffusion Policy. Experiments on five challenging real-world tasks show that TacThru-UMI achieves an average success rate of 85.5%, significantly outperforming the baselines of tactile policy(66.3%) and vision-only policy (55.4%). The system excels in critical scenarios, including contact detection with thin and soft objects and precision manipulation requiring multimodal coordination. This work demonstrates that combining simultaneous multimodal perception with modern learning frameworks enables more precise, adaptable robotic manipulation.

多模态感知机器人抓取触觉传感模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。