arXiv:2509.06155cs.CV2025-09被引 62

统一生成音视频,用专家拼接提升效率与同步性。

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

  • 通过拼接预训练音视频专家模型,避免从零训练
  • 7600小时数据微调后,环境音与语音生成高度同步
  • 自研在线标注流程,解决文本标签对齐问题

我们提出UniVerse-1,一个类似Veo-3的统一音视频生成模型,可同时生成协调的音频与视频。为提升训练效率,采用专家拼接(SoE)技术,深度融合预训练视频与音乐生成模型的对应模块,充分复用其基础能力。为确保环境音与语音在时间上与视频内容精准对齐,开发了在线标注流水线,在训练过程中实时生成标签,避免因文本标注错位导致性能下降。经约7600小时音视频数据微调后,模型在环境音生成中实现良好视听协同,在语音生成中表现出强对齐性。为系统评估方法,我们构建新基准数据集Verse-Bench。为推动音视频生成研究并缩小与Veo3等顶尖模型的差距,已公开模型与代码。项目页:https://dorniwang.github.io/UniVerse-1/

原文摘要 · Abstract (English)

We introduce UniVerse-1, a unified, Veo-3-like model capable of simultaneously generating coordinated audio and video. To enhance training efficiency, we bypass training from scratch and instead employ a stitching of experts (SoE) technique. This approach deeply fuses the corresponding blocks of pre-trained video and music generation experts models, thereby fully leveraging their foundational capabilities. To ensure accurate annotations and temporal alignment for both ambient sounds and speech with video content, we developed an online annotation pipeline that processes the required training data and generates labels during training process. This strategy circumvents the performance degradation often caused by misalignment text-based annotations. Through the synergy of these techniques, our model, after being finetuned on approximately 7,600 hours of audio-video data, produces results with well-coordinated audio-visuals for ambient sounds generation and strong alignment for speech generation. To systematically evaluate our proposed method, we introduce Verse-Bench, a new benchmark dataset. In an effort to advance research in audio-video generation and to close the performance gap with state-of-the-art models such as Veo3, we make our model and code publicly available. We hope this contribution will benefit the broader research community. Project page: https://dorniwang.github.io/UniVerse-1/.

音视频生成专家拼接多模态同步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。