Ovi统一生成音视频,实现自然同步,无需分步处理。
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
- 用双塔DiT结构块级融合音视频信息,实时对齐时间与语义。
- 在数十万小时原始音频上训练,生成逼真语音与环境音效。
- 适合影视生成、虚拟主播等需要音画协同的场景。
音视频生成常依赖复杂多阶段架构或顺序合成。本文提出Ovi,一种将音视频视为单一生成过程的统一范式。通过双塔DiT模块的块级跨模态融合,Ovi实现自然同步,无需独立流水线或后期对齐。为实现细粒度多模态融合,音频塔采用与强预训练视频模型相同的架构,并在数十万小时原始音频上从零训练,学习生成逼真音效及富含说话人身份与情感的语音。融合通过在大规模视频数据集上联合训练视频与音频塔实现,利用缩放RoPE嵌入传递时间信息,双向交叉注意力传递语义信息。模型可生成具备自然语音与精准上下文匹配音效的电影级视频片段。所有演示、代码与模型权重已公开于https://aaxwaz.github.io/Ovi。
原文摘要 · Abstract (English)
Audio-video generation has often relied on complex multi-stage architectures or sequential synthesis of sound and visuals. We introduce Ovi, a unified paradigm for audio-video generation that models the two modalities as a single generative process. By using blockwise cross-modal fusion of twin-DiT modules, Ovi achieves natural synchronization and removes the need for separate pipelines or post hoc alignment. To facilitate fine-grained multimodal fusion modeling, we initialize an audio tower with an architecture identical to that of a strong pretrained video model. Trained from scratch on hundreds of thousands of hours of raw audio, the audio tower learns to generate realistic sound effects, as well as speech that conveys rich speaker identity and emotion. Fusion is obtained by jointly training the identical video and audio towers via blockwise exchange of timing (via scaled-RoPE embeddings) and semantics (through bidirectional cross-attention) on a vast video corpus. Our model enables cinematic storytelling with natural speech and accurate, context-matched sound effects, producing movie-grade video clips. All the demos, code and model weights are published at https://aaxwaz.github.io/Ovi
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。