通过解耦语义特征空间,实现高保真人脸动画与精准动作迁移
Learning Semantic Facial Descriptors for Accurate Face Animation
- 构建可学习的解耦向量空间,分离身份与动作特征
- 在VoxCeleb、HDTF、CelebV上优于现有方法,身份保留更佳
- 适合需要高质量人脸动画的生成与迁移任务
人脸动画是一项挑战性任务。基于模型的方法(如3DMM或关键点)常导致类模型重建效果,难以有效保持身份特征;而无模型方法则面临难以获得解耦且语义丰富的特征空间,从而难以实现精确的动作迁移。为此,本文引入可学习的语义人脸描述符,在解耦的向量空间中同时赋予身份与动作子空间语义。通过在源脸和驱动脸上使用编码器获取正交基向量系数,得到身份与动作子空间中的有效描述符。最终将这些描述符重新组合为潜在代码以实现人脸动画。在VoxCeleb、HDTF、CelebV三个挑战性基准上的大量实验表明,本方法在身份保留和动作迁移方面均超越当前最优方法,取得更优的定量与定性结果。
原文摘要 · Abstract (English)
Face animation is a challenging task. Existing model-based methods (utilizing 3DMMs or landmarks) often result in a model-like reconstruction effect, which doesn't effectively preserve identity. Conversely, model-free approaches face challenges in attaining a decoupled and semantically rich feature space, thereby making accurate motion transfer difficult to achieve. We introduce the semantic facial descriptors in learnable disentangled vector space to address the dilemma. The approach involves decoupling the facial space into identity and motion subspaces while endowing each of them with semantics by learning complete orthogonal basis vectors. We obtain basis vector coefficients by employing an encoder on the source and driving faces, leading to effective facial descriptors in the identity and motion subspaces. Ultimately, these descriptors can be recombined as latent codes to animate faces. Our approach successfully addresses the issue of model-based methods' limitations in high-fidelity identity and the challenges faced by model-free methods in accurate motion transfer. Extensive experiments are conducted on three challenging benchmarks (i.e. VoxCeleb, HDTF, CelebV). Comprehensive quantitative and qualitative results demonstrate that our model outperforms SOTA methods with superior identity preservation and motion transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。