arXiv:2512.02005cs.CV2025-12

用声音识别物体可交互区域,让模型听声识物。

Learning Visual Affordance from Audio

  • 用音频与视觉融合建模,实现声音驱动的交互区域分割
  • 在自建数据集上达到当前最佳性能,零样本泛化效果显著
  • 适合做多模态理解、具身智能和机器人交互的研究者

我们提出音频-视觉可交互性定位(AV-AG)新任务,通过动作声音分割物体的可交互区域。与依赖文本指令或示范视频的方法不同,音频提供实时、语义丰富且与视觉无关的线索,更直观地揭示交互区域。为此,我们构建了首个AV-AG数据集,包含大量动作音频、物体图像及像素级可交互标注,还设置了未见子集用于评估零样本泛化能力。同时提出AVAGFormer模型,采用语义条件交叉模态混合器和双头解码器,有效融合音视频信号进行掩码预测。实验表明,AVAGFormer在该任务上优于相关基准方法。综合分析揭示了AV-AG与视觉分割(AVS)的本质区别,验证了端到端建模的优势及各模块贡献。代码与数据集已开源。

原文摘要 · Abstract (English)

We introduce Audio-Visual Affordance Grounding (AV-AG), a new task that segments object interaction regions from action sounds. Unlike existing approaches that rely on textual instructions or demonstration videos, which often limited by ambiguity or occlusion, audio provides real-time, semantically rich, and visually independent cues for affordance grounding, enabling more intuitive understanding of interaction regions. To support this task, we construct the first AV-AG dataset, comprising a large collection of action sounds, object images, and pixel-level affordance annotations. The dataset also includes an unseen subset to evaluate zero-shot generalization. Furthermore, we propose AVAGFormer, a model equipped with a semantic-conditioned cross-modal mixer and a dual-head decoder that effectively fuses audio and visual signals for mask prediction. Experiments show that AVAGFormer achieves state-of-the-art performance on AV-AG, surpassing baselines from related tasks. Comprehensive analyses highlight the distinctions between AV-AG and AVS, the benefits of end-to-end modeling, and the contribution of each component. Code and dataset have been released on https://jscslld.github.io/AVAGFormer/.

多模态音频感知视觉定位具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。