arXiv:2601.19606cs.CVcs.AI2026-01

多尺度对齐与生成预训练提升音视频对应建模效果

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining

  • 设计多尺度对比学习,捕捉不同粒度的音视频语义与时间关联
  • 引入时空扩散生成目标,实现音视频跨模态转换与合成
  • 在VGGSound等数据集上超越现有方法,适合音视频生成与检索任务

近期视频-音频(V-A)理解与生成进展越来越多地依赖于联合的音视频嵌入,为跨模态检索和生成任务提供基础。尽管先前方法如CAVP通过对比目标有效建模了模态间的语义与时间对应关系,但其性能仍不理想。主要瓶颈在于未能充分建模视频与音频信号的密集多尺度特性,对应关系常跨越从细粒度到粗粒度的空间-时间结构,而现有框架对此利用不足。为此,我们提出GMS-CAVP,一种结合多尺度音视频对齐与多尺度时空扩散式预训练目标的新框架,以增强音视频对应建模。首先,GMS-CAVP引入多尺度对比学习策略,捕捉不同粒度下的语义与时间关系。其次,我们突破传统对比学习,引入基于扩散的生成目标,实现音视频之间的模态转换与合成。这种统一的判别-生成范式促进更深层次的跨模态理解,并为高保真生成铺平道路。在VGGSound、AudioSet和Panda70M上的大量实验表明,GMS-CAVP在生成与检索任务中均优于以往方法。

原文摘要 · Abstract (English)

Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retrieval and generation. While prior methods like CAVP effectively model semantic and temporal correspondences between modalities using contrastive objectives, their performance remains suboptimal. A key limitation is the insufficient modeling of the dense, multi-scale nature of both video and audio signals, correspondences often span fine- to coarse-grained spatial-temporal structures, which are underutilized in existing frameworks. To this end, we propose GMS-CAVP, a novel framework that combines Multi-Scale Video-Audio Alignment and Multi-Scale Spatial-Temporal Diffusion-based pretraining objectives to enhance V-A correspondence modeling. First, GMS-CAVP introduces a multi-scale contrastive learning strategy that captures semantic and temporal relations across varying granularities. Second, we go beyond traditional contrastive learning by incorporating a diffusion-based generative objective, enabling modality translation and synthesis between video and audio. This unified discriminative-generative formulation facilitates deeper cross-modal understanding and paves the way for high-fidelity generation. Extensive experiments on VGGSound, AudioSet, and Panda70M demonstrate that GMS-CAVP outperforms previous methods in generation and retrieval.

音视频对齐多尺度建模扩散模型生成预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。