arXiv:2511.12072cs.MMcs.AI2025-11

用投影潜空间提升音视频生成效率与同步性

ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation

  • 将音频转为视频样表示,统一时空维度对齐模态
  • 多尺度注意力机制实现细粒度时序建模与跨模态融合
  • 3D潜空间扩散变压器降低计算开销,适合实时生成

由于音频与视频在结构上存在固有错位,且多模态数据处理计算成本高,音视频生成(SVG)仍具挑战。本文提出ProAV-DiT,一种用于高效同步音视频生成的投影潜空间扩散变换器。为解决结构不一致问题,我们预先将原始音频转换为视频样表征,对齐其时间与空间维度。核心采用多尺度双流时空自编码器(MDSA),通过正交分解将两模态投影至统一潜空间,实现细粒度时空建模与语义对齐。为进一步增强时序连贯性与模态特异性融合,引入多尺度注意力机制,包含多尺度时间自注意力与分组跨模态注意力。此外,将MDSA输出的2D潜变量堆叠为统一3D潜空间,由时空扩散变换器处理,有效建模时空依赖关系,生成高质量同步音视频内容的同时降低计算开销。在标准基准上的大量实验表明,ProAV-DiT在生成质量与计算效率方面均优于现有方法。

原文摘要 · Abstract (English)

Sounding Video Generation (SVG) remains a challenging task due to the inherent structural misalignment between audio and video, as well as the high computational cost of multimodal data processing. In this paper, we introduce ProAV-DiT, a Projected Latent Diffusion Transformer designed for efficient and synchronized audio-video generation. To address structural inconsistencies, we preprocess raw audio into video-like representations, aligning both the temporal and spatial dimensions between audio and video. At its core, ProAV-DiT adopts a Multi-scale Dual-stream Spatio-Temporal Autoencoder (MDSA), which projects both modalities into a unified latent space using orthogonal decomposition, enabling fine-grained spatiotemporal modeling and semantic alignment. To further enhance temporal coherence and modality-specific fusion, we introduce a multi-scale attention mechanism, which consists of multi-scale temporal self-attention and group cross-modal attention. Furthermore, we stack the 2D latents from MDSA into a unified 3D latent space, which is processed by a spatio-temporal diffusion Transformer. This design efficiently models spatiotemporal dependencies, enabling the generation of high-fidelity synchronized audio-video content while reducing computational overhead. Extensive experiments conducted on standard benchmarks demonstrate that ProAV-DiT outperforms existing methods in both generation quality and computational efficiency.

音视频生成扩散模型多模态对齐高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。