arXiv:2409.14608cs.RO2024-09被引 4

用声音和视觉结合检测机器人抓握时的外部接触点。

Visual-auditory Extrinsic Contact Estimation

  • 用麦克风和扬声器主动发声,通过声音传播感知外部接触
  • 仿真训练后零样本迁移到真实场景,准确定位接触位置与范围
  • 适合需要精细触觉反馈的复杂抓取任务,如避障或堆叠

可靠的抓握操作依赖于机器人对物体与环境之间外部接触的感知。然而,由于遮挡、分辨率有限及近接触状态模糊,仅靠视觉难以有效观测此类接触。本文提出一种视觉-听觉融合方法,结合全局视觉场景信息与主动音频传感获取的局部接触信号。系统在机械夹爪上集成接触式麦克风与导电扬声器,通过被抓物体传递声波,检测外部接触。整个感知流程在仿真中训练,并实现零样本迁移至真实世界。为弥合仿真到现实的差距,引入真实音频幻觉技术,将真实音频样本注入带有真实接触标签的仿真场景。所提出的多模态模型在多种杂乱与遮挡场景中,均能准确估计外部接触的位置与尺寸。进一步实验表明,显式接触预测显著提升了下游接触密集型操作任务中的策略学习效果。

原文摘要 · Abstract (English)

Robust manipulation often hinges on a robot's ability to perceive extrinsic contacts-contacts between a grasped object and its surrounding environment. However, these contacts are difficult to observe through vision alone due to occlusions, limited resolution, and ambiguous near-contact states. In this paper, we propose a visual-auditory method for extrinsic contact estimation that integrates global scene information from vision with local contact cues obtained through active audio sensing. Our approach equips a robotic gripper with contact microphones and conduction speakers, enabling the system to emit and receive acoustic signals through the grasped object to detect external contacts. We train our perception pipeline entirely in simulation and zero-shot transfer to the real world. To bridge the sim-to-real gap, we introduce a real-to-sim audio hallucination technique, injecting real-world audio samples into simulated scenes with ground-truth contact labels. The resulting multimodal model accurately estimates both the location and size of extrinsic contacts across a range of cluttered and occluded scenarios. We further demonstrate that explicit contact prediction significantly improves policy learning for downstream contact-rich manipulation tasks.

触觉感知多模态机器人操作听觉传感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。