MoCoTalk通过自适应路由实现多条件控制,生成更逼真的说话头视频。
MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation

- 引入自适应路由机制,动态融合参考图像、关键点、光照网格和语音四类条件。
- 在多个指标上达到顶尖表现,唇部动作与语音对齐更精准。
- 支持属性级控制,适合需要精细调节的视频生成场景。
说话头生成需联合建模身份、头部姿态、面部表情和嘴部动态。现有方法通常仅处理部分因素,且在多条件融合时依赖固定权重或启发式策略。本文提出MoCoTalk,一种统一四类互补控制信号的视频扩散框架:参考图像、面部关键点、3DMM渲染的阴影网格及对应语音音频。为缓解异构条件间的破坏性干扰,设计了自适应多条件路由,计算通道级、时间感知的门控,使融合策略随特征子空间和噪声水平动态变化。为更好捕捉语音相关的面部动态,提出唇部增强阴影网格,一种基于3DMM的表示,可解耦头部运动、嘴部运动、表情与光照。该设计提供时序一致的几何先验,并允许推理时灵活重组属性。此外引入唇部一致性损失,强化音画对齐。大量实验表明,MoCoTalk在多数结构、运动与感知指标上达当前最优,同时提供单条件方法无法实现的属性级可控性。
原文摘要 · Abstract (English)
Talking-head generation requires joint modeling of identity, head pose, facial expression, and mouth dynamics. Existing methods typically address only a subset of these factors, and rely on fixed-weight or heuristic fusion when multiple conditions are involved. We present MoCoTalk, a multi-conditional video diffusion framework that unifies four complementary control signals: a reference image, facial keypoints, 3DMM-rendered shading meshes, and the corresponding speech audio. To resolve destructive interference among heterogeneous conditions, we introduce an Adaptive Multi-Condition Router that computes channel-wise, timestep-aware gating over the four condition streams, allowing the fusion strategy to vary with both feature subspace and noise level. To better capture speech-related facial dynamics, we design a Mouth-Augmented Shading Mesh, a 3DMM-based representation that decouples head motion, mouth motion, expression, and lighting. This design provides a temporally consistent geometric prior and allows flexible recombination of these attributes at inference. We further introduce a lip consistency loss to tighten audio-visual alignment. Extensive experiments show that MoCoTalk achieves state-of-the-art performance on the majority of structural, motion, and perceptual metrics, while offering attribute-level controllability that single-condition methods do not provide.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。