arXiv:2603.16093cs.SDcs.AI2026-03

提出首个端到端音视频联合生成框架,支持高质量同步输出。

Diffusion Models for Joint Audio-Video Generation

  • 构建13小时游戏+64小时演出的配对数据集,每段34秒便于复现。
  • 新架构在快速动作与音乐线索上实现精准对齐,生成质量高。
  • 分步生成法先出视频再补音轨,适配多模态应用开发者。

多模态生成模型在单模态音视频合成方面取得显著进展,但真正的音视频联合生成仍是开放挑战。本文提出四项关键贡献:首先,发布两个高质量配对音视频数据集,包含13小时游戏片段和64小时音乐会表演,均分割为一致的34秒样本,以促进可复现研究;其次,在这些数据集上从头训练MM-Diffusion架构,证明其能生成语义一致的音视频对,并在快速动作和音乐线索上进行定量对齐评估;第三,探究基于预训练视频与音频编码器-解码器的联合隐空间扩散,揭示多模态解码阶段存在挑战与不一致性;最后,提出一种两阶段文本到音视频生成流水线:先生成视频,再以视频输出和原始提示为条件生成时间同步的音频。实验表明,该模块化方法实现了高质量音视频联合生成。

原文摘要 · Abstract (English)

Multimodal generative models have shown remarkable progress in single-modality video and audio synthesis, yet truly joint audio-video generation remains an open challenge. In this paper, I explore four key contributions to advance this field. First, I release two high-quality, paired audio-video datasets. The datasets consisting on 13 hours of video-game clips and 64 hours of concert performances, each segmented into consistent 34-second samples to facilitate reproducible research. Second, I train the MM-Diffusion architecture from scratch on our datasets, demonstrating its ability to produce semantically coherent audio-video pairs and quantitatively evaluating alignment on rapid actions and musical cues. Third, I investigate joint latent diffusion by leveraging pretrained video and audio encoder-decoders, uncovering challenges and inconsistencies in the multimodal decoding stage. Finally, I propose a sequential two-step text-to-audio-video generation pipeline: first generating video, then conditioning on both the video output and the original prompt to synthesize temporally synchronized audio. My experiments show that this modular approach yields high-fidelity generations of audio video generation.

音视频生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。