提出分层融合与表征对齐,提升音视频分离效果
Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
- 分层融合:结合中间与晚期融合优势,适应不同声音特性
- 在MUSIC等数据集上达到最新水平,超越现有自监督方法
- 通过预训练音频嵌入对齐,缩小音视频模态差异
自监督音视频源分离利用音频与视觉模态间的自然关联来分离混合音频信号。本文系统分析了现有多模态融合方法在音视频分离任务中的表现,发现中间融合更适合处理短时瞬态声音,而晚期融合更擅长捕捉持续且谐波丰富的声音。为此,我们提出一种分层融合策略,有效整合两个融合阶段。此外,通过引入高质量外部音频表征而非仅依赖音频分支独立学习,可使训练更简单。为此,我们提出表征对齐方法,将音频编码器的隐层特征与预训练音频模型提取的嵌入对齐。在MUSIC、MUSIC-21和VGGSound数据集上的大量实验表明,该方法在自监督设置下达到当前最优性能。进一步分析显示,表征对齐有效缩小了音频与视觉模态间的差距。
原文摘要 · Abstract (English)
Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematically analyse the performance of existing multimodal fusion methods for audio-visual separation task, demonstrating that the performance of different fusion strategies is closely linked to the characteristics of the sound: middle fusion is better suited for handling short, transient sounds, while late fusion is more effective for capturing sustained and harmonically rich sounds. We thus propose a hierarchical fusion strategy that effectively integrates both fusion stages. In addition, training can be made easier by incorporating high-quality external audio representations, rather than relying solely on the audio branch to learn them independently. To explore this, we propose a representation alignment approach that aligns the latent features of the audio encoder with embeddings extracted from pre-trained audio models. Extensive experiments on MUSIC, MUSIC-21 and VGGSound datasets demonstrate that our approach achieves state-of-the-art results, surpassing existing methods under the self-supervised setting. We further analyse the impact of representation alignment on audio features, showing that it reduces modality gap between the audio and visual modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。