arXiv:2512.13677cs.CV2025-12

统一视频音频生成与编辑,用联合表示提升多模态交互效率

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

  • 双分支架构实现视频、音频、文本原生联合建模
  • 在多个数据集上达到当前最佳性能,支持生成与编辑任务
  • 适合内容创作、多媒体生成领域研究者参考

本文提出JoVA,一个统一的视频-音频联合生成与编辑框架。现有方法常依赖碎片化、任务特定结构或复杂融合机制,而JoVA采用原生联合表征学习,在双分支架构中实现视频、音频与文本的直接交互,消除冗余对齐模块,有效整合多种多模态任务于单一模型。此外,通过通道条件控制实现灵活图像与视频参考,避免大量标记膨胀,并引入嘴部区域损失以增强口型对齐。为充分赋能并系统评估该框架,我们构建了涵盖视频-音频生成与编辑的数据集,设计了针对性的统一基准。大量实验表明,JoVA在多个基准上均取得领先性能,验证了其作为多功能内容创作框架的可扩展性。

原文摘要 · Abstract (English)

In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on fragmented, task-specific architectures or complex fusion mechanisms, JoVA employs native joint representation learning for direct video, audio, and text interaction in a dual-branch architecture. This design eliminates redundant alignment modules and effectively unifies diverse multimodal tasks within a single model. Furthermore, we utilize channel-wise conditioning for flexible image and video reference to avoid massive token expansion, alongside a mouth-area loss to enhance lip alignment. To fully empower and systematically evaluate this framework, we construct a comprehensive training corpus encompassing video-audio generation and editing datasets, and introduce unified benchmarks tailored for these multimodal tasks. Extensive experiments demonstrate that JoVA achieves state-of-the-art performance across benchmarks, establishing it as an extensible framework for versatile content creation. Project page: https://visual-ai.github.io/jova

多模态生成视频音频联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。