arXiv:2602.19163cs.CVcs.MM2026-02中稿 · ICLR被引 18

JavisDiT++统一建模与优化,实现高质量音视频同步生成。

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

论文配图:JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
图 1 · 摘自论文原文
  • 采用模态专用专家混合架构,提升跨模态交互与单模生成质量。
  • 通过帧级时序对齐位置编码,实现音视频在时间上的精确同步。
  • 基于人类偏好优化,显著提升生成内容的流畅性与一致性。

AIGC已从文本到图像生成拓展至高质量多模态音视频合成。在此背景下,联合音视频生成(JAVG)作为关键任务,能从文本描述中生成同步且语义一致的音视频内容。然而,相较于Veo3等先进商用模型,现有开源方法在生成质量、时间同步性和人类偏好对齐方面仍存在不足。为此,本文提出JavisDiT++,一个简洁而强大的统一建模范式与优化框架。首先,引入模态专用专家混合(MS-MoE)设计,提升跨模态交互效率并增强单模生成质量;其次,提出时序对齐旋转位置编码(TA-RoPE),实现音频与视频标记间的显式帧级同步;此外,开发音视频直接偏好优化(AV-DPO)方法,使模型输出在质量、一致性和同步性维度上更符合人类偏好。基于Wan2.1-1.3B-T2V,仅用约100万条公开训练数据即达到当前最佳性能,在定性与定量评估中均优于先前方法。全面消融实验验证了各模块有效性。代码、模型与数据集均已开源:https://JavisVerse.github.io/JavisDiT2-page。

原文摘要 · Abstract (English)

AIGC has rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision from textual descriptions. However, compared with advanced commercial models such as Veo3, existing open-source methods still suffer from limitations in generation quality, temporal synchrony, and alignment with human preferences. To bridge the gap, this paper presents JavisDiT++, a concise yet powerful framework for unified modeling and optimization of JAVG. First, we introduce a modality-specific mixture-of-experts (MS-MoE) design that enables cross-modal interaction efficacy while enhancing single-modal generation quality. Then, we propose a temporal-aligned RoPE (TA-RoPE) strategy to achieve explicit, frame-level synchronization between audio and video tokens. Besides, we develop an audio-video direct preference optimization (AV-DPO) method to align model outputs with human preference across quality, consistency, and synchrony dimensions. Built upon Wan2.1-1.3B-T2V, our model achieves state-of-the-art performance merely with around 1M public training entries, significantly outperforming prior approaches in both qualitative and quantitative evaluations. Comprehensive ablation studies have been conducted to validate the effectiveness of our proposed modules. All the code, model, and dataset are released at https://JavisVerse.github.io/JavisDiT2-page.

音视频生成多模态扩散模型偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。