arXiv:2511.07619cs.RO2025-11被引 2

机器人通过听觉视觉探索,高效学习物体的音视特性。

CAVER: Curious Audiovisual Exploring Robot

  • 用3D打印末端执行器激发物体发声,获取音视数据。
  • 结合视觉与声音特征,构建全局局部融合的表示。
  • 基于好奇心驱动探索,减少交互次数实现高效覆盖。

多模态音视感知可为机器人操作开辟新路径,如提升材料分类能力或仅凭音频模仿示范(如听音演奏)。但要释放这种潜力,机器人需学习物体外观与其互动发声之间的关联。这要求具备新交互能力、表征方式和探索策略,以高效积累丰富的音视知识。本文提出CAVER,一种能构建并利用丰富音视表征的新型机器人,包含三项创新:1)可安装于平行夹爪的3D打印新型末端执行器,用于激发物体声学响应;2)融合局部与全局视觉信息及声音特征的音视表征;3)基于好奇心驱动的探索算法,优先与高不确定性物体交互,以较少交互次数获得对意外声音的良好覆盖。实验表明,CAVER在多种场景下比多个探索基线更高效地构建表征,并显著提升材料分类性能与音频示范模仿效果。

原文摘要 · Abstract (English)

Multimodal audiovisual perception can enable new avenues for robotic manipulation, from better material classification to the imitation of demonstrations for which only audio signals are available (e.g., playing a tune by ear). However, to unlock such multimodal potential, robots need to learn the correlations between an object's visual appearance and the sound it generates when they interact with it. Such an active sensorimotor experience requires new interaction capabilities, representations, and exploration methods to guide the robot in efficiently building increasingly rich audiovisual knowledge. In this work, we present CAVER, a novel robot that builds and utilizes rich audiovisual representations of objects. CAVER includes three novel contributions: 1) a novel 3D printed end-effector, attachable to parallel grippers, that excites objects' audio responses, 2) an audiovisual representation that combines local and global appearance information with sound features, and 3) an exploration algorithm that uses and builds the audiovisual representation in a curiosity-driven manner that prioritizes interacting with high uncertainty objects to obtain good coverage of surprising audio with fewer interactions. We demonstrate that CAVER builds rich representations in different scenarios more efficiently than several exploration baselines, and that the learned audiovisual representation leads to significant improvements in material classification and the imitation of audio-only human demonstrations. https://caver-bot.github.io/

机器人音视融合好奇心探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。