arXiv:2505.01237cs.MMcs.CV2025-05CVPR被引 17

通过细粒度对齐提升音视频自监督学习效果

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

  • 将音频视为与视频帧对齐的时序序列,实现跨模态精细对齐
  • 分离对比学习与重建目标,用专用全局标记缓解优化冲突
  • 引入可学习注册标记,提升空间定位精度,适合多模态研究者

近期音视频学习进展在跨模态表示学习方面取得显著成果。然而,现有方法大多依赖全局音频表示,难以捕捉与视觉帧的细粒度时间对应关系。同时,联合学习重建与跨模态对齐常面临优化目标冲突问题。本文提出 CAV-MAE Sync,作为原始 CAV-MAE 框架的简单但有效的扩展,用于自监督音视频学习。针对三大挑战:首先,通过将音频建模为与视频帧对齐的时序序列,解决模态间粒度不匹配问题;其次,通过专用全局标记分离对比学习与重建目标,缓解优化冲突;第三,引入可学习注册标记,减轻补丁标记的语义负担,提升空间定位能力。在 AudioSet、VGG Sound 与 ADE20K Sound 数据集上的零样本检索、分类与定位任务中,该方法表现优异,达到当前最优水平,优于更复杂架构。

原文摘要 · Abstract (English)

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences with visual frames. Additionally, existing methods often struggle with conflicting optimization objectives when trying to jointly learn reconstruction and cross-modal alignment. In this work, we propose CAV-MAE Sync as a simple yet effective extension of the original CAV-MAE framework for self-supervised audio-visual learning. We address three key challenges: First, we tackle the granularity mismatch between modalities by treating audio as a temporal sequence aligned with video frames, rather than using global representations. Second, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens. Third, we improve spatial localization by introducing learnable register tokens that reduce semantic load on patch tokens. We evaluate the proposed approach on AudioSet, VGG Sound, and the ADE20K Sound dataset on zero-shot retrieval, classification and localization tasks demonstrating state-of-the-art performance and outperforming more complex architectures.

音视频对齐自监督学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。