arXiv:2602.13640cs.ROcs.AI2026-02

用声音信号提升机器人抓取精度,解决视觉看不清时的触觉感知难题。

Hierarchical Audio-Visual-Proprioceptive Fusion for Precise Robotic Manipulation

  • 分层融合声音、视觉和本体感觉,先用声音引导其他模态
  • 在倒液体、开柜门任务中表现优于现有方法
  • 适合需要精准触觉反馈的复杂物理操作场景

现有机器人操作方法主要依赖视觉与本体感觉,但在部分可观测的真实环境中难以推断接触状态。声音信号天然蕴含丰富的交互动态,却在多模态融合中未被充分利用。当前多数融合方法假设各模态角色相同,采用扁平对称结构,但此假设不适用于本质稀疏且由接触驱动的声音信号。为实现基于声学感知的精确操作,我们提出一种分层表示融合框架,逐步整合音频、视觉与本体感觉。该方法首先以声音信号条件化视觉与本体表示,再显式建模高阶跨模态交互,捕捉模态间的互补依赖关系。融合后的表示用于基于扩散模型的策略,直接从多模态观测生成连续机器人动作。端到端学习与分层结构使策略能有效利用任务相关的声学信息,同时抑制低信息量模态的干扰。方法在真实世界任务(如液体倾倒、柜门开启)上评估,结果表明其持续优于最先进多模态融合框架,尤其在声音提供视觉无法获取的任务相关信息时。此外,通过互信息分析揭示了声音在多模态融合中的作用机制。

原文摘要 · Abstract (English)

Existing robotic manipulation methods primarily rely on visual and proprioceptive observations, which may struggle to infer contact-related interaction states in partially observable real-world environments. Acoustic cues, by contrast, naturally encode rich interaction dynamics during contact, yet remain underexploited in current multimodal fusion literature. Most multimodal fusion approaches implicitly assume homogeneous roles across modalities, and thus design flat and symmetric fusion structures. However, this assumption is ill-suited for acoustic signals, which are inherently sparse and contact-driven. To achieve precise robotic manipulation through acoustic-informed perception, we propose a hierarchical representation fusion framework that progressively integrates audio, vision, and proprioception. Our approach first conditions visual and proprioceptive representations on acoustic cues, and then explicitly models higher-order cross-modal interactions to capture complementary dependencies among modalities. The fused representation is leveraged by a diffusion-based policy to directly generate continuous robot actions from multimodal observations. The combination of end-to-end learning and hierarchical fusion structure enables the policy to exploit task-relevant acoustic information while mitigating interference from less informative modalities. The proposed method has been evaluated on real-world robotic manipulation tasks, including liquid pouring and cabinet opening. Extensive experiment results demonstrate that our approach consistently outperforms state-of-the-art multimodal fusion frameworks, particularly in scenarios where acoustic cues provide task-relevant information not readily available from visual observations alone. Furthermore, a mutual information analysis is conducted to interpret the effect of audio cues in robotic manipulation via multimodal fusion.

多模态融合机器人操作声音感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。