arXiv:2509.19999cs.MMcs.CV2025-09

多事件视频生成音频,通过双流对比预训练与偏好优化提升精准度

MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization

  • 采用双流对比音频视觉预训练,同步对齐语义与动态特征
  • 引入直接偏好优化,显著提升音频质量与时空对齐精度
  • 适合复杂多声源场景的音频生成研究者与应用开发者

当前视频到音频(V2A)方法在包含多个音源、声事件或过渡的复杂多事件场景中表现不佳,主要受限于两个核心问题:一是难以精确对齐复杂的语义信息与快速的动态特征;二是基础训练缺乏对语义-时序对齐和音频质量的量化偏好优化。为此,本文提出新型V2A框架MultiSoundGen,首次将直接偏好优化(DPO)引入V2A领域,并结合音频-视觉预训练(AVP)提升复杂场景下的生成能力。核心创新包括:1)提出SlowFast Contrastive AVP(SF-CAVP),一种统一双流架构的AVP模型,显式对齐视听数据的核心语义表示与快速动态特征;2)提出基于AVP的偏好优化(AVP-RPO),利用SF-CAVP作为奖励模型,量化并优先优化关键语义-时序匹配,同时提升音频质量。实验表明,MultiSoundGen在多事件场景下达到当前最优性能,在分布匹配、音频质量、语义对齐与时序同步方面均取得全面提升。

原文摘要 · Abstract (English)

Current video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical limitations. First, existing methods face challenges in precisely aligning intricate semantic information together with rapid dynamic features. Second, foundational training lacks quantitative preference optimization for semantic-temporal alignment and audio quality. As a result, it fails to enhance integrated generation quality in cluttered multi-event scenes. To address these core limitations, this study proposes a novel V2A framework: MultiSoundGen. It introduces direct preference optimization (DPO) into the V2A domain, leveraging audio-visual pretraining (AVP) to enhance performance in complex multi-event scenarios. Our contributions include two key innovations: the first is SlowFast Contrastive AVP (SF-CAVP), a pioneering AVP model with a unified dual-stream architecture. SF-CAVP explicitly aligns core semantic representations and rapid dynamic features of audio-visual data to handle multi-event complexity; second, we integrate the DPO method into V2A task and propose AVP-Ranked Preference Optimization (AVP-RPO). It uses SF-CAVP as a reward model to quantify and prioritize critical semantic-temporal matches while enhancing audio quality. Experiments demonstrate that MultiSoundGen achieves state-of-the-art (SOTA) performance in multi-event scenarios, delivering comprehensive gains across distribution matching, audio quality, semantic alignment, and temporal synchronization. Demos are available at https://v2aresearch.github.io/MultiSoundGen/.

视频生成音频生成多事件偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。