arXiv:2410.01594cs.CV2024-10中稿 · ACM MM 2024被引 20

用统一图像表示音频视频,实现高质量音视频联合生成

MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation

  • 将音视频转为图像统一表示,构建分层多模态编码器
  • 在多个数据集上提升生成质量与速度,超越现有方法
  • 支持长视频生成、音频延续等任务,通用性强

音视频联合生成(SVG)面临高维信号空间、不同数据格式和内容模式差异的挑战。为此,我们提出一种新型多模态潜变量扩散模型(MM-LDM)。首先将音频与视频数据转换为单一或一对图像,实现统一表示;随后引入分层多模态自编码器,分别为每种模态构建低层感知潜空间,以及共享的高层语义特征空间。前者在感知上等价于原始信号空间,但大幅降低维度;后者弥合模态间信息鸿沟,提供更深入的跨模态引导。所提方法在所有评估指标上均取得显著提升,并在Landscape与AIST++数据集上实现更快训练与采样速度。进一步评估显示,其在开放域音视频生成、长时音视频生成、音频延续、视频延续及条件单模态生成任务中展现出优异的适应性与泛化能力。

原文摘要 · Abstract (English)

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content information. To address these issues, we introduce a novel multi-modal latent diffusion model (MM-LDM) for the SVG task. We first unify the representation of audio and video data by converting them into a single or a couple of images. Then, we introduce a hierarchical multi-modal autoencoder that constructs a low-level perceptual latent space for each modality and a shared high-level semantic feature space. The former space is perceptually equivalent to the raw signal space of each modality but drastically reduces signal dimensions. The latter space serves to bridge the information gap between modalities and provides more insightful cross-modal guidance. Our proposed method achieves new state-of-the-art results with significant quality and efficiency gains. Specifically, our method achieves a comprehensive improvement on all evaluation metrics and a faster training and sampling speed on Landscape and AIST++ datasets. Moreover, we explore its performance on open-domain sounding video generation, long sounding video generation, audio continuation, video continuation, and conditional single-modal generation tasks for a comprehensive evaluation, where our MM-LDM demonstrates exciting adaptability and generalization ability.

音视频生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。