arXiv:2608.16285cs.CVcs.AI2026-08

用深度信息提升音视频分割精度,让模型更懂物体空间位置。

Audio-Visual Segmentation via Depth-Guided Collaborative Modeling

论文配图:Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
图 1 · 摘自论文原文
  • 引入估计深度作为空间结构线索,三模态协同建模音视频与深度信息。
  • 在AVSS数据集上,平均交并比提升10.2%,对复杂遮挡场景效果显著。
  • 适合做音视频理解、自动驾驶等需要精准目标定位的场景研究。

音视频分割(AVS)是多模态感知的基础任务,通过融合视觉与音频线索实现视频中发声物体的像素级分割,在视频理解、人机交互和自动驾驶中有广泛应用。然而,现有方法通常未显式建模相对距离、遮挡等几何线索,限制了跨模态对齐的鲁棒性。人类感知中,空间结构自然融入音视频证据以精确定位发声物体。受此启发,本文引入估计深度作为空间结构线索,提出DGCM-AVS三模态框架,联合建模音频、视觉与深度信息。设计深度感知动态调制模块,增强相邻物体分离能力并保持物体内特征一致性;提出深度引导渐进融合机制,利用深度作为中间桥梁逐步对齐音频与视觉特征。相比最先进方法,DGCM-AVS在AVSS数据集上取得M_J提升10.2%、M_F提升8.7%的性能,验证了深度作为关键但被忽视的模态在AVS中的潜力,为后续研究提供新方向。

原文摘要 · Abstract (English)

Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues. It has broad applications in video understanding, human-computer interaction, and autonomous driving. However, most existing AVS methods do not explicitly model geometric cues such as relative distance and occlusion, thereby limiting the robustness of cross-modal alignment. In human perception, spatial structure is naturally integrated with audio-visual evidence to accurately localize sounding objects. Motivated by this, we incorporate estimated depth as a spatial structural cue for AVS and propose DGCM-AVS, a tri-modal framework that jointly models audio, visual, and depth information. Specifically, we design a Depth-Aware Dynamic Modulator to improve the separation of adjacent objects while preserving intra-object feature consistency. Furthermore, we propose Depth-Guided Progressive Fusion, which uses depth as an intermediate bridge to progressively align audio cues with visual features. Compared to state-of-the-art methods, DGCM-AVS achieves relative improvements of 10.2 percent in M_J and 8.7 percent in M_F on the AVSS dataset. We believe our study highlights depth as a promising yet underexplored modality for AVS and may encourage further research in this direction.

音视频分割深度引导多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。