arXiv:2603.18600cs.CV2026-03被引 2

提升音视频联合生成质量,解决跨模态对齐与训练不一致问题。

Improving Joint Audio-Video Generation with Cross-Modal Context Learning

  • 设计跨模态上下文学习框架,优化音视频特征对齐与信息融合。
  • 在少数据条件下实现顶尖生成效果,资源消耗显著降低。
  • 适合音视频生成、多模态模型研究者参考,尤其关注训练稳定性。

基于双流Transformer的音视频联合生成方法已成为当前主流。通过结合预训练的视频扩散模型和音频扩散模型,并引入跨模态交互注意力模块,可在极少训练数据下生成高质量、时序同步的音视频内容。本文重新审视该范式,分析其局限性:门控机制导致的模型流形变化、跨模态注意力引入的多模态背景区域偏差、训练与推理中多模态分类器无关引导(CFG)的不一致性,以及多重条件间的冲突。为缓解上述问题,提出跨模态上下文学习(CCL)框架,包含多个精心设计模块:时序对齐的RoPE与分块(TARP)有效增强音频隐变量与视频隐变量的时序对齐;跨模态上下文注意力(CCA)中的可学习上下文标记(LCT)与动态上下文路由(DCR)提供稳定无条件锚点,按任务动态路由,加速收敛并提升生成质量;推理阶段的无条件上下文引导(UCG)利用LCT提供的无条件支持,适配多种CFG形式,改善训练-推理一致性,缓解冲突。全面评估表明,CCL在性能上达到最新水平,同时大幅降低资源需求。

原文摘要 · Abstract (English)

The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a cross-modal interaction attention module, high-quality, temporally synchronized audio-video content can be generated with minimal training data. In this paper, we first revisit the dual-stream transformer paradigm and further analyze its limitations, including model manifold variations caused by the gating mechanism controlling cross-modal interactions, biases in multi-modal background regions introduced by cross-modal attention, and the inconsistencies in multi-modal classifier-free guidance (CFG) during training and inference, as well as conflicts between multiple conditions. To alleviate these issues, we propose Cross-Modal Context Learning (CCL), equipped with several carefully designed modules. Temporally Aligned RoPE and Partitioning (TARP) effectively enhances the temporal alignment between audio latent and video latent representations. The Learnable Context Tokens (LCT) and Dynamic Context Routing (DCR) in the Cross-Modal Context Attention (CCA) module provide stable unconditional anchors for cross-modal information, while dynamically routing based on different training tasks, further enhancing the model's convergence speed and generation quality. During inference, Unconditional Context Guidance (UCG) leverages the unconditional support provided by LCT to facilitate different forms of CFG, improving train-inference consistency and further alleviating conflicts. Through comprehensive evaluations, CCL achieves state-of-the-art performance compared with recent academic methods while requiring substantially fewer resources.

音视频生成跨模态扩散模型生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。