arXiv:2511.03334cs.CV2025-11被引 34

统一音视频生成框架,提升同步与语义一致性。

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

  • 双分支扩散模型+非对称跨模态交互,实现精准时空对齐。
  • 仅用130万样本训练,同步与情感一致性优于3010万样本模型。
  • 适合音视频联合生成、配音替换等多任务场景。

由于缺乏有效的跨模态建模,现有开源音视频生成方法常出现口型不同步和语义不一致的问题。为此,我们提出UniAVGen,一个统一的音视频联合生成框架。该框架基于双分支联合合成架构,采用两个并行的扩散Transformer(DiTs)构建连贯的跨模态潜在空间。其核心是非对称跨模态交互机制,支持双向、时序对齐的跨注意力,确保精确的时空同步与语义一致性。此外,通过面部感知调制模块动态强化关键区域的交互。为提升推理阶段生成质量,引入模态感知无分类器引导策略,显式增强跨模态相关性信号。值得一提的是,UniAVGen的强健联合生成设计可无缝统一多项关键音视频任务,包括联合生成与延续、视频转音频配音、音频驱动视频合成。大量实验表明,仅需130万训练样本(对比3010万),其在音视频同步性、音色一致性和情绪一致性方面均表现更优。

原文摘要 · Abstract (English)

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a unified framework for joint audio and video generation. UniAVGen is anchored in a dual-branch joint synthesis architecture, incorporating two parallel Diffusion Transformers (DiTs) to build a cohesive cross-modal latent space. At its heart lies an Asymmetric Cross-Modal Interaction mechanism, which enables bidirectional, temporally aligned cross-attention, thus ensuring precise spatiotemporal synchronization and semantic consistency. Furthermore, this cross-modal interaction is augmented by a Face-Aware Modulation module, which dynamically prioritizes salient regions in the interaction process. To enhance generative fidelity during inference, we additionally introduce Modality-Aware Classifier-Free Guidance, a novel strategy that explicitly amplifies cross-modal correlation signals. Notably, UniAVGen's robust joint synthesis design enables seamless unification of pivotal audio-video tasks within a single model, such as joint audio-video generation and continuation, video-to-audio dubbing, and audio-driven video synthesis. Comprehensive experiments validate that, with far fewer training samples (1.3M vs. 30.1M), UniAVGen delivers overall advantages in audio-video synchronization, timbre consistency, and emotion consistency.

音视频生成扩散模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。