JavisDiT实现音视频同步生成,提升真实场景下的内容质量与对齐精度。
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

- 基于分层时空先验同步机制,实现音视频精细对齐
- 在10,140个高质量音视频数据上验证,同步误差显著降低
- 适合需要高保真音视频协同生成的研究与应用
本文提出JavisDiT,一种用于音视频同步生成(JAVG)的联合音频-视频扩散变换器。基于强大的扩散变换器(DiT)架构,JavisDiT在统一框架中从开放文本提示同时生成高质量音视频内容。为确保音视频同步,引入细粒度时空对齐机制,通过分层时空同步先验(HiST-Sypo)估计器提取全局与细粒度时空先验,引导视觉与听觉组件的同步。此外,构建新基准JavisBench,包含10,140个高质量带文字描述的发声视频,聚焦复杂真实场景中的同步评估。并设计鲁棒指标,量化真实内容中生成音视频对的同步性。实验表明,JavisDiT在保证高质量生成的同时,显著优于现有方法,确立了JAVG任务的新标准。代码、模型与数据见https://javisverse.github.io/JavisDiT-page/。
原文摘要 · Abstract (English)
This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates high-quality audio and video content from open-ended user prompts in a unified framework. To ensure audio-video synchronization, we introduce a fine-grained spatio-temporal alignment mechanism through a Hierarchical Spatial-Temporal Synchronized Prior (HiST-Sypo) Estimator. This module extracts both global and fine-grained spatio-temporal priors, guiding the synchronization between the visual and auditory components. Furthermore, we propose a new benchmark, JavisBench, which consists of 10,140 high-quality text-captioned sounding videos and focuses on synchronization evaluation in diverse and complex real-world scenarios. Further, we specifically devise a robust metric for measuring the synchrony between generated audio-video pairs in real-world content. Experimental results demonstrate that JavisDiT significantly outperforms existing methods by ensuring both high-quality generation and precise synchronization, setting a new standard for JAVG tasks. Our code, model, and data are available at https://javisverse.github.io/JavisDiT-page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。