arXiv:2509.19749cs.CV2025-09被引 1

用面部动作单元精准控制说话头生成表情,更自然真实。

Talking Head Generation via AU-Guided Landmark Prediction

  • 通过显式建模动作单元到面部关键点的映射,实现逐帧表情控制。
  • 在MEAD数据集上多指标优于当前最佳方法,表达更准确稳定。
  • 适合需要精细表情控制的虚拟人、影视制作场景。

我们提出一种两阶段音视频驱动说话头生成框架,通过面部动作单元(AUs)实现细粒度表情控制。与依赖情感标签或隐式AU条件的方法不同,本模型显式将AUs映射到2D面部关键点,实现物理合理的逐帧表达控制。第一阶段采用变分运动生成器,从音频和AU强度预测时序连贯的关键点序列;第二阶段使用基于扩散的合成器,根据这些关键点和参考图像生成逼真且口型同步的视频。该运动与外观分离的设计提升了表达准确性、时间稳定性与视觉真实感。在MEAD数据集上的实验表明,本方法在多个指标上超越现有最优基线,验证了显式AU-to-landmark建模在表达性说话头生成中的有效性。

原文摘要 · Abstract (English)

We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotion labels or implicit AU conditioning, our model explicitly maps AUs to 2D facial landmarks, enabling physically grounded, per-frame expression control. In the first stage, a variational motion generator predicts temporally coherent landmark sequences from audio and AU intensities. In the second stage, a diffusion-based synthesizer generates realistic, lip-synced videos conditioned on these landmarks and a reference image. This separation of motion and appearance improves expression accuracy, temporal stability, and visual realism. Experiments on the MEAD dataset show that our method outperforms state-of-the-art baselines across multiple metrics, demonstrating the effectiveness of explicit AU-to-landmark modeling for expressive talking head generation.

说话头生成动作单元关键点控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。