分离高阶语义与低阶细节建模,提升说话头像的音视频同步质量。
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

- 用共享骨干+专用解码器,分层处理音视频特征
- 在基准测试中音视频同步率、画质、音质均优于双分支模型
- 适合需要高质量音视频协同生成的数字人应用
联合音视频生成模型表明,统一生成比串行方法具有更强的跨模态一致性。然而,现有模型通过广泛注意力在去噪过程中始终耦合模态,将高层语义与底层细节完全纠缠,这在说话头像合成中并不理想:尽管音频与面部动作在语义上相关,但其底层实现(声学信号与视觉纹理)遵循不同的渲染过程。强制所有层级联合建模导致不必要的纠缠并降低效率。我们提出 Talker-T2AV,一种自回归扩散框架,其中高层跨模态建模在共享主干中完成,而底层细化使用模态特定解码器。一个共享的自回归语言模型在统一的补丁级标记空间中联合推理音频与视频。两个轻量级扩散Transformer头将隐藏状态解码为帧级音频与视频潜在表示。在说话头像基准测试中,Talker-T2AV 在唇同步准确率、视频质量与音频质量上均优于双分支基线模型,展现出强于串行流水线的跨模态一致性。
原文摘要 · Abstract (English)
Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating high-level semantics and low-level details in a fully entangled manner. This is suboptimal for talking head synthesis: while audio and facial motion are semantically correlated, their low-level realizations (acoustic signals and visual textures) follow distinct rendering processes. Enforcing joint modeling across all levels causes unnecessary entanglement and reduces efficiency. We propose Talker-T2AV, an autoregressive diffusion framework where high-level cross-modal modeling occurs in a shared backbone, while low-level refinement uses modality-specific decoders. A shared autoregressive language model jointly reasons over audio and video in a unified patch-level token space. Two lightweight diffusion transformer heads decode the hidden states into frame-level audio and video latents. Experiments on talking portrait benchmarks show Talker-T2AV outperforms dual-branch baselines in lip-sync accuracy, video quality, and audio quality, achieving stronger cross-modal consistency than cascaded pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。