arXiv:2508.13442cs.CV2025-08被引 6

实现人脸动作全解耦,支持音视频驱动的可控说话头生成

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis

  • 将嘴型、姿态、眼神、表情分离到四个独立潜空间,用可学习基底线性组合表示
  • 各空间基底正交,训练时自动分配运动责任,实现无干扰解耦控制
  • 支持音频/视频输入,适合需要精细控制的虚拟主播与影视特效应用

实现多面部动作的解耦控制并兼容多样输入模态,显著提升说话头生成的应用与娱乐价值。现有方法常忽视面部特征解耦空间的深度探索,导致各动作间相互干扰且难以共享。为此,本文提出EDTalk++,一种全新的全解耦可控说话头生成框架。该框架可独立操控嘴型、头部姿态、眼动与情感表达,适配视频或音频输入。具体地,采用四个轻量模块将面部动态分解为对应嘴型、姿态、眼神和表情的四类潜空间,每空间由一组可学习基底构成,其线性组合定义特定动作。通过强制基底间正交并设计高效训练策略,确保各空间独立且无需外部知识即可自动分配运动职责。学习到的基底存储于对应库中,实现与音频输入共享视觉先验。此外,针对各空间特性,提出音频转动作模块,用于音频驱动的说话头生成。实验验证了EDTalk++的有效性。

原文摘要 · Abstract (English)

Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation. This necessitates a deep exploration of the decoupling space for facial features, ensuring that they a) operate independently without mutual interference and b) can be preserved to share with different modal inputs, both aspects often neglected in existing methods. To address this gap, this paper proposes EDTalk++, a novel full disentanglement framework for controllable talking head generation. Our framework enables individual manipulation of mouth shape, head pose, eye movement, and emotional expression, conditioned on video or audio inputs. Specifically, we employ four lightweight modules to decompose the facial dynamics into four distinct latent spaces representing mouth, pose, eye, and expression, respectively. Each space is characterized by a set of learnable bases whose linear combinations define specific motions. To ensure independence and accelerate training, we enforce orthogonality among bases and devise an efficient training strategy to allocate motion responsibilities to each space without relying on external knowledge. The learned bases are then stored in corresponding banks, enabling shared visual priors with audio input. Furthermore, considering the properties of each space, we propose an Audio-to-Motion module for audio-driven talking head synthesis. Experiments are conducted to demonstrate the effectiveness of EDTalk++.

说话头生成解耦控制音频驱动潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。