arXiv:2503.12840cs.SDcs.CV2025-03CVPR被引 12

解决声音混杂与匹配困难,提升音视频分割精度

Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics

  • 通过动态提取与消除机制,分离并筛选音频特征
  • 在AVS数据集上实现优于现有方法的分割性能
  • 适合研究多模态感知与音视频对齐的学者

声源引导的目标分割因其增强多模态感知的潜力而受到广泛关注。以往方法主要聚焦于设计先进架构以促进有效的音视频交互,但未充分应对音频本身的固有挑战:(1)音频信号重叠导致的特征混淆;(2)同一物体产生不同声音带来的音视频匹配困难。为此,我们提出动态提取与消除框架DDESeg:为缓解特征混淆,DDESeg通过增强各声源的独特语义信息,重构混合音频的语义内容,提取保留各自特性的表示;为降低匹配难度,引入判别性特征学习模块,增强生成音频表示的语义区分度;考虑到并非所有提取的音频表示都对应视觉特征(如非可视区域的声音),提出动态消除模块,过滤不匹配元素,促进发声区域与相关音频语义的精准交互。通过评分交互特征,识别并剔除无关音频信息,确保准确的音视频对齐。大量实验证明,本框架在多个音视频分割数据集上表现优异。

原文摘要 · Abstract (English)

Sound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions, without fully addressing the inherent challenges posed by audio natures, \emph{\ie}, (1) feature confusion due to the overlapping nature of audio signals, and (2) audio-visual matching difficulty from the varied sounds produced by the same object. To address these challenges, we propose Dynamic Derivation and Elimination (DDESeg): a novel audio-visual segmentation framework. Specifically, to mitigate feature confusion, DDESeg reconstructs the semantic content of the mixed audio signal by enriching the distinct semantic information of each individual source, deriving representations that preserve the unique characteristics of each sound. To reduce the matching difficulty, we introduce a discriminative feature learning module, which enhances the semantic distinctiveness of generated audio representations. Considering that not all derived audio representations directly correspond to visual features (e.g., off-screen sounds), we propose a dynamic elimination module to filter out non-matching elements. This module facilitates targeted interaction between sounding regions and relevant audio semantics. By scoring the interacted features, we identify and filter out irrelevant audio information, ensuring accurate audio-visual alignment. Comprehensive experiments demonstrate that our framework achieves superior performance in AVS datasets.

音视频分割多模态音频增强语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。