arXiv:2503.23377cs.CVcs.AI2025-03中稿 · ICLR被引 75

JavisDiT实现音视频同步生成,提升真实场景下的内容质量与对齐精度。

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

论文配图:JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
图 1 · 摘自论文原文
  • 基于分层时空先验同步机制,实现音视频精细对齐
  • 在10,140个高质量音视频数据上验证,同步误差显著降低
  • 适合需要高保真音视频协同生成的研究与应用

本文提出JavisDiT,一种用于音视频同步生成(JAVG)的联合音频-视频扩散变换器。基于强大的扩散变换器(DiT)架构,JavisDiT在统一框架中从开放文本提示同时生成高质量音视频内容。为确保音视频同步,引入细粒度时空对齐机制,通过分层时空同步先验(HiST-Sypo)估计器提取全局与细粒度时空先验,引导视觉与听觉组件的同步。此外,构建新基准JavisBench,包含10,140个高质量带文字描述的发声视频,聚焦复杂真实场景中的同步评估。并设计鲁棒指标,量化真实内容中生成音视频对的同步性。实验表明,JavisDiT在保证高质量生成的同时,显著优于现有方法,确立了JAVG任务的新标准。代码、模型与数据见https://javisverse.github.io/JavisDiT-page/。

原文摘要 · Abstract (English)

This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates high-quality audio and video content from open-ended user prompts in a unified framework. To ensure audio-video synchronization, we introduce a fine-grained spatio-temporal alignment mechanism through a Hierarchical Spatial-Temporal Synchronized Prior (HiST-Sypo) Estimator. This module extracts both global and fine-grained spatio-temporal priors, guiding the synchronization between the visual and auditory components. Furthermore, we propose a new benchmark, JavisBench, which consists of 10,140 high-quality text-captioned sounding videos and focuses on synchronization evaluation in diverse and complex real-world scenarios. Further, we specifically devise a robust metric for measuring the synchrony between generated audio-video pairs in real-world content. Experimental results demonstrate that JavisDiT significantly outperforms existing methods by ensuring both high-quality generation and precise synchronization, setting a new standard for JAVG tasks. Our code, model, and data are available at https://javisverse.github.io/JavisDiT-page/.

音视频生成扩散模型同步对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。