让机器人通过声音定位和识别物体,实现更智能的抓取操作。
S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

- 融合视觉与声学空间信息,构建多模态模仿学习框架。
- 在需同时判断位置和音色的任务中表现最优,实测有效。
- 适合对声音感知有要求的现实机械臂操控场景。
声音能提供关于物体位置、材质属性以及接触或运动引发变化的丰富线索。本文提出一套新的声学感知操控任务,要求机器人利用听觉线索确定操作目标。这些任务需结合声音源定位与识别,以实现机器人操控中的主动探索。为此,我们提出一种多模态模仿学习框架——空间-频谱音频动作(S2A2),将视觉特征与声学空间及信号信息融合,用于声学感知操控任务。我们实现了S2A2模型,整合了ACT、Diffusion Policy、VQ-BeT和$π_0$等策略。仿真实验表明,该方法在需同时判断位置与音色的任务中效果最佳;真实机器人实验进一步验证了所提任务与框架在真实世界操控中的可行性。
原文摘要 · Abstract (English)
Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and $π_0$, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。