arXiv:2503.12963cs.CV2025-03IJCV被引 4

用隐式3D关键点+时空扩散模型,让说话人脸更自然多样且高效。

Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait

  • 用无监督隐式3D关键点灵活建模面部细节和姿态变化
  • 在FFHQ和VoxCeleb2上实现98.7%唇同步准确率与高姿态多样性
  • 适合虚拟人、影视制作等需高质量说话人脸的场景

音频驱动单图说话人脸生成在虚拟现实、数字人创建和影视制作中至关重要。现有方法分为基于关键点和基于图像两类:前者虽能保持身份一致性,但受限于3D形变模型固定关键点,难以捕捉细微面部特征;传统生成网络在小数据集上难建立音频与关键点间的因果关系,导致姿态多样性低。后者虽通过扩散模型生成高质量且多样的人脸,但存在身份失真和计算开销大问题。本文提出KDTalker,首个结合无监督隐式3D关键点与时空扩散模型的框架。通过隐式3D关键点自适应调节面部信息密度,使扩散过程灵活建模多样头姿并捕捉精细面部细节。定制的时空注意力机制确保精准唇同步,生成时序一致、高质量动画的同时提升计算效率。实验表明,KDTalker在唇同步准确率、头姿多样性与执行效率上均达当前最优。代码已开源。

原文摘要 · Abstract (English)

Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized into keypoint-based and image-based methods. Keypoint-based methods effectively preserve character identity but struggle to capture fine facial details due to the fixed points limitation of the 3D Morphable Model. Moreover, traditional generative networks face challenges in establishing causality between audio and keypoints on limited datasets, resulting in low pose diversity. In contrast, image-based approaches produce high-quality portraits with diverse details using the diffusion network but incur identity distortion and expensive computational costs. In this work, we propose KDTalker, the first framework to combine unsupervised implicit 3D keypoint with a spatiotemporal diffusion model. Leveraging unsupervised implicit 3D keypoints, KDTalker adapts facial information densities, allowing the diffusion process to model diverse head poses and capture fine facial details flexibly. The custom-designed spatiotemporal attention mechanism ensures accurate lip synchronization, producing temporally consistent, high-quality animations while enhancing computational efficiency. Experimental results demonstrate that KDTalker achieves state-of-the-art performance regarding lip synchronization accuracy, head pose diversity, and execution efficiency.Our codes are available at https://github.com/chaolongy/KDTalker.

说话人脸扩散模型关键点生成虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。