UniTalking统一生成逼真说话人视频,支持语音克隆与精准口型同步。
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation

- 用多模态变换器建模音视频时序对齐,共享自注意力机制提升同步精度。
- 在公开数据集上口型同步准确率超现有开源方法,视觉质量达顶尖水平。
- 支持仅用短音频参考克隆目标语音风格,适合个性化视频生成应用。
尽管当前最先进的音视频生成模型如Veo3和Sora2表现出色,但其闭源特性导致架构与训练方式不可访问。为弥合这一可及性与性能差距,我们提出UniTalking——一种统一的端到端扩散框架,用于生成高保真语音与唇形同步视频。核心采用多模态变换器块,通过共享自注意力机制显式建模音频与视频隐变量之间的细粒度时序对应关系。利用预训练视频生成模型的强大先验知识,该框架在保证顶尖视觉质量的同时实现高效训练。此外,UniTalking集成个性化语音克隆能力,可基于简短音频参考生成目标风格语音。定性和定量结果表明,该方法生成的说话人视频高度逼真,在唇形同步准确率、语音自然度及整体感知质量方面均优于现有开源方案。
原文摘要 · Abstract (English)
While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and performance, we introduce UniTalking, a unified, end-to-end diffusion framework for generating high-fidelity speech and lip-synchronized video. At its core, our framework employs Multi-Modal Transformer Blocks to explicitly model the fine-grained temporal correspondence between audio and video latent tokens via a shared self-attention mechanism. By leveraging powerful priors from a pre-trained video generation model, our framework ensures state-of-the-art visual fidelity while enabling efficient training. Furthermore, UniTalking incorporates a personalized voice cloning capability, allowing the generation of speech in a target style from a brief audio reference. Qualitative and quantitative results demonstrate that our method produces highly realistic talking portraits, achieving superior performance over existing open-source approaches in lip-sync accuracy, audio naturalness, and overall perceptual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。