arXiv:2508.12163cs.CVcs.AI2025-08ICCV

让虚拟人脸说话时表情自然可控,还能保持原样貌。

RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis

  • 用音频生成3D面部关键点,结合情绪标签调表情
  • 在多个评测中情绪准确率超现有方法15%以上
  • 适合需要真实情感交互的虚拟主播、AI助手

情绪是人工社交智能的核心。当前方法虽在口型同步和图像质量上表现优异,但在生成准确且可控制的情绪表达方面仍存在不足,且难以保持主体身份特征。为此,我们提出 RealTalk,一种能够实现高情绪准确率、强情绪可控性与鲁棒身份保留的逼真情感化说话头合成框架。RealTalk 利用变分自编码器(VAE)从驱动音频生成3D面部关键点,通过基于ResNet的地标变形模型(LDM),将关键点与情绪标签嵌入向量拼接,生成带有情绪特征的关键点;这些关键点与面部混合形状系数共同作为条件,输入新型三平面注意力神经辐射场(NeRF)以合成高度逼真的情感化说话头。大量实验表明,RealTalk 在情绪准确率、可控性与身份保留方面均优于现有方法,推动了社会智能AI系统的发展。

原文摘要 · Abstract (English)

Emotion is a critical component of artificial social intelligence. However, while current methods excel in lip synchronization and image quality, they often fail to generate accurate and controllable emotional expressions while preserving the subject's identity. To address this challenge, we introduce RealTalk, a novel framework for synthesizing emotional talking heads with high emotion accuracy, enhanced emotion controllability, and robust identity preservation. RealTalk employs a variational autoencoder (VAE) to generate 3D facial landmarks from driving audio, which are concatenated with emotion-label embeddings using a ResNet-based landmark deformation model (LDM) to produce emotional landmarks. These landmarks and facial blendshape coefficients jointly condition a novel tri-plane attention Neural Radiance Field (NeRF) to synthesize highly realistic emotional talking heads. Extensive experiments demonstrate that RealTalk outperforms existing methods in emotion accuracy, controllability, and identity preservation, advancing the development of socially intelligent AI systems.

人脸生成情绪识别NeRF语音驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。