arXiv:2510.02313cs.CV2025-10ICCV

模型学会区分物体碰撞声,如勺子敲地板与地毯的声音差异。

Clink! Chop! Thud! -- Learning Object Sounds from Real-World Interactions

  • 用第一视角视频和自动分割掩码引导模型聚焦物体交互区域。
  • 在新任务上达到当前最优性能,同时提升其他多模态动作理解效果。
  • 适合对声音感知、物体交互建模感兴趣的视觉-听觉研究者。

日常物体互动会产生独特的声响。我们提出“发声物体检测”任务,评估模型将声音与直接参与的物体关联的能力。受人类感知启发,我们的多模态物体感知框架从真实场景的第一人称视频中学习。为促进以物体为中心的方法,我们首先设计自动管道,计算交互物体的分割掩码,指导模型训练时关注最信息丰富的区域。采用槽注意力视觉编码器进一步强化物体先验。我们在新任务及现有多模态动作理解任务上均取得领先性能。

原文摘要 · Abstract (English)

Can a model distinguish between the sound of a spoon hitting a hardwood floor versus a carpeted one? Everyday object interactions produce sounds unique to the objects involved. We introduce the sounding object detection task to evaluate a model's ability to link these sounds to the objects directly involved. Inspired by human perception, our multimodal object-aware framework learns from in-the-wild egocentric videos. To encourage an object-centric approach, we first develop an automatic pipeline to compute segmentation masks of the objects involved to guide the model's focus during training towards the most informative regions of the interaction. A slot attention visual encoder is used to further enforce an object prior. We demonstrate state of the art performance on our new task along with existing multimodal action understanding tasks.

声音识别多模态物体检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。