arXiv:2510.24103cs.SDcs.AI2025-10NeurIPS被引 1

自引导双角色对齐,让视频生成音频更真实

Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

  • 用自引导机制让模型自己优化音视频对齐
  • 在VGGSound上FAD降至0.40,超越现有方法
  • 适合做跨模态生成、音视频合成的研究者

我们提出MGAudio,一种面向开放域视频转音频的新型基于流的框架,核心是模型引导的双角色对齐设计。不同于依赖分类器或无分类器引导的先前方法,MGAudio通过专为视频条件音频生成设计的训练目标,使生成模型能够自我引导。该框架包含三个主要组件:(1) 可扩展的基于流的Transformer模型,(2) 双角色对齐机制,其中音视频编码器同时作为条件输入模块和特征对齐器以提升生成质量,(3) 模型引导目标,增强跨模态一致性和音频真实感。MGAudio在VGGSound数据集上取得最先进性能,FAD降低至0.40,显著优于最佳无分类器引导基线,并在FD、IS和对齐度指标上持续领先。同时在挑战性UnAV-100基准上表现出良好泛化能力。结果表明,模型引导的双角色对齐是条件视频到音频生成中一个强大且可扩展的新范式。代码已公开于:https://github.com/pantheon5100/mgaudio

原文摘要 · Abstract (English)

We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The framework integrates three main components: (1) a scalable flow-based Transformer model, (2) a dual-role alignment mechanism where the audio-visual encoder serves both as a conditioning module and as a feature aligner to improve generation quality, and (3) a model-guided objective that enhances cross-modal coherence and audio realism. MGAudio achieves state-of-the-art performance on VGGSound, reducing FAD to 0.40, substantially surpassing the best classifier-free guidance baselines, and consistently outperforms existing methods across FD, IS, and alignment metrics. It also generalizes well to the challenging UnAV-100 benchmark. These results highlight model-guided dual-role alignment as a powerful and scalable paradigm for conditional video-to-audio generation. Code is available at: https://github.com/pantheon5100/mgaudio

视频生成音视频对齐流模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。