实现零样本语音克隆的多模态音视频联合生成模型
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
- 用统一指令-音素输入实现音视频精准对齐
- 零样本语音克隆效果媲美专业语音合成系统
- 适合需要高保真音视频生成的研究与应用
联合音视频生成旨在合成同步的多感官内容,但现有统一模型在细粒度声学控制方面表现不佳,尤其在保持语音身份的前提下。现有方法或因级联生成导致时间错位,或无法在联合生成框架内实现零样本语音克隆。本文提出MM-Sonate,一种统一的多模态流匹配框架,兼具可控音视频联合生成与零样本语音克隆能力。不同于依赖粗略语义描述的方法,MM-Sonate采用统一指令-音素输入以确保严格的语言与时间对齐。为实现零样本语音克隆,引入音色注入机制,有效分离说话人身份与语言内容。针对标准无分类器引导在多模态设置中的局限性,提出基于噪声的负条件策略,利用自然噪声先验显著提升声学保真度。实证评估表明,MM-Sonate在联合生成基准上达到新SOTA,显著优于基线,在唇形同步和语音可懂度方面表现突出,同时语音克隆保真度接近专用文本转语音系统。
原文摘要 · Abstract (English)
Joint audio-video generation aims to synthesize synchronized multisensory content, yet current unified models struggle with fine-grained acoustic control, particularly for identity-preserving speech. Existing approaches either suffer from temporal misalignment due to cascaded generation or lack the capability to perform zero-shot voice cloning within a joint synthesis framework. In this work, we present MM-Sonate, a multimodal flow-matching framework that unifies controllable audio-video joint generation with zero-shot voice cloning capabilities. Unlike prior works that rely on coarse semantic descriptions, MM-Sonate utilizes a unified instruction-phoneme input to enforce strict linguistic and temporal alignment. To enable zero-shot voice cloning, we introduce a timbre injection mechanism that effectively decouples speaker identity from linguistic content. Furthermore, addressing the limitations of standard classifier-free guidance in multimodal settings, we propose a noise-based negative conditioning strategy that utilizes natural noise priors to significantly enhance acoustic fidelity. Empirical evaluations demonstrate that MM-Sonate establishes new state-of-the-art performance in joint generation benchmarks, significantly outperforming baselines in lip synchronization and speech intelligibility, while achieving voice cloning fidelity comparable to specialized Text-to-Speech systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。