根据声音定位视频中碰撞物体,提升人机感知协同能力。
Segmenting Collision Sound Sources in Egocentric Videos
- 结合音频与视觉线索,用基础模型实现弱监督分割。
- 在两个新数据集上mIoU提升3倍至4.7倍,显著优于基线。
- 适合研究视听感知、人机交互与主动感知的学者参考。
人类擅长多感官感知,常能通过物体交互声音识别其属性。受此启发,我们提出全新任务——碰撞声音源分割(CS3),旨在根据音频信息,在视觉输入(即碰撞视频帧)中分割出引发碰撞的声音物体。该任务面临独特挑战:碰撞声源于两物体交互,声学特征同时依赖两者。研究聚焦于第一人称视频,其中声音清晰但场景杂乱、物体小且互动短暂。为此,我们提出一种弱监督音频条件分割方法,利用基础模型(CLIP与SAM2),并引入第一人称视角线索(如手部物体)以定位可能的碰撞源。在自建的两个基准数据集EPIC-CS3与Ego4D-CS3上,本方法在mIoU指标上分别超越基线3倍与4.7倍。
原文摘要 · Abstract (English)
Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sound Source Segmentation (CS3), where we aim to segment the objects responsible for a collision sound in visual input (i.e. video frames from the collision clip), conditioned on the audio. This task presents unique challenges. Unlike isolated sound events, a collision sound arises from interactions between two objects, and the acoustic signature of the collision depends on both. We focus on egocentric video, where sounds are often clear, but the visual scene is cluttered, objects are small, and interactions are brief. To address these challenges, we propose a weakly-supervised method for audio-conditioned segmentation, utilising foundation models (CLIP and SAM2). We also incorporate egocentric cues, i.e. objects in hands, to find acting objects that can potentially be collision sound sources. Our approach outperforms competitive baselines by $3\times$ and $4.7\times$ in mIoU on two benchmarks we introduce for the CS3 task: EPIC-CS3 and Ego4D-CS3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。