arXiv:2503.19001cs.CVcs.AI2025-03被引 2

用解耦扩散模型实现跨语言人脸动画,兼顾精准控制与时间连贯性。

DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model

  • 将3DMM表达参数解耦为语义子空间,实现区域级精细控制
  • 在3DMM参数空间构建分层扩散架构,提升唇形同步与表情质量
  • 自建中文高清数据集CHDTF,适配中文说话人脸生成任务

近年来,说话人脸生成技术取得了显著进展。然而,现有方法存在根本性局限:基于3DMM的方法虽能保持时间一致性,但缺乏细粒度区域控制;基于Stable Diffusion的方法可实现空间操作,却存在时间不一致问题。两者的融合受限于控制机制不兼容及面部表征的语义纠缠。本文提出DisentTalk,引入数据驱动的语义解耦框架,将3DMM表达参数分解为有意义的子空间,实现细粒度面部控制。在此解耦表示基础上,构建在3DMM参数空间运行的分层潜在扩散架构,并引入区域感知注意力机制,确保空间精度与时间连贯性。为缓解高质量中文训练数据稀缺问题,我们构建了CHDTF——一个中文高清说话人脸数据集。大量实验表明,该方法在多个指标上优于现有方法,包括唇形同步、表情质量和时间一致性。项目主页:https://kangweiiliu.github.io/DisentTalk。

原文摘要 · Abstract (English)

Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional control, while Stable Diffusion-based methods enable spatial manipulation but suffer from temporal inconsistencies. The integration of these approaches is hindered by incompatible control mechanisms and semantic entanglement of facial representations. This paper presents DisentTalk, introducing a data-driven semantic disentanglement framework that decomposes 3DMM expression parameters into meaningful subspaces for fine-grained facial control. Building upon this disentangled representation, we develop a hierarchical latent diffusion architecture that operates in 3DMM parameter space, integrating region-aware attention mechanisms to ensure both spatial precision and temporal coherence. To address the scarcity of high-quality Chinese training data, we introduce CHDTF, a Chinese high-definition talking face dataset. Extensive experiments show superior performance over existing methods across multiple metrics, including lip synchronization, expression quality, and temporal consistency. Project Page: https://kangweiiliu.github.io/DisentTalk.

说话人脸扩散模型跨语言解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。