arXiv:2412.15220cs.MMcs.SD2024-12被引 22

同步生成音视频,让声音与画面完全对齐。

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

  • 用双扩散变换器架构联合建模音视频,实现信息融合。
  • 生成的音视频相关性更强,音频质量显著提升。
  • 支持零样本音视频转换和新分辨率适配,无需重训练。

视频与音频是人类自然感知的紧密关联模态。尽管近期研究已能从文本生成音频或视频,但同时生成两者仍多依赖级联流程或多模态对比编码器,常因推理与条件化过程中的信息丢失导致效果不佳。本文提出SyncFlow,可从文本同步生成时序对齐的音视频。其核心为双扩散-变压器(d-DiT)架构,实现音视频联合建模与有效信息融合。为降低联合建模的计算开销,SyncFlow采用分阶段训练策略:先分别学习音视频,再进行联合微调。实证评估表明,SyncFlow生成的音视频在相关性上优于基线方法,音频质量显著提升,且具有强零样本能力,包括零样本视频转音频生成及无需再训练即可适应新视频分辨率。

原文摘要 · Abstract (English)

Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on either a cascaded process or multi-modal contrastive encoders. These approaches, however, often lead to suboptimal results due to inherent information losses during inference and conditioning. In this paper, we introduce SyncFlow, a system that is capable of simultaneously generating temporally synchronized audio and video from text. The core of SyncFlow is the proposed dual-diffusion-transformer (d-DiT) architecture, which enables joint video and audio modelling with proper information fusion. To efficiently manage the computational cost of joint audio and video modelling, SyncFlow utilizes a multi-stage training strategy that separates video and audio learning before joint fine-tuning. Our empirical evaluations demonstrate that SyncFlow produces audio and video outputs that are more correlated than baseline methods with significantly enhanced audio quality and audio-visual correspondence. Moreover, we demonstrate strong zero-shot capabilities of SyncFlow, including zero-shot video-to-audio generation and adaptation to novel video resolutions without further training.

音视频生成扩散模型多模态同步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。