让虚拟人脸随语音自然表达情绪,效果远超现有方法。
EmoGene: Audio-Driven Emotional 3D Talking-Head Generation
- 用音频生成面部动作点,再融合情绪特征生成情感化表情点
- 在多个数据集上实现更高画质与更准确的情绪同步
- 适合虚拟人交互、影视制作等需要真实情感表达的场景
音频驱动的虚拟人脸生成是虚拟人互动与影视制作中的关键技术。尽管近期研究多聚焦于图像质量和口型同步,但情绪表达的生成仍缺乏深入探索。本文提出 EmoGene,一种新型高保真、音频驱动的情感化人脸视频生成框架。该方法采用基于变分自编码器(VAE)的音视频到动作模块生成面部关键点,并在动作到情绪模块中将其与情绪嵌入融合,输出情感关键点;这些关键点驱动基于神经辐射场(NeRF)的情绪到视频模块,生成逼真的情感化说话头视频。此外,我们设计了一种姿态采样方法,用于在无语音输入时生成自然的静止状态视频。大量实验表明,EmoGene 在生成高保真情感说话头视频方面显著优于现有方法。
原文摘要 · Abstract (English)
Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional expressions remains underexplored. In this paper, we introduce EmoGene, a novel framework for synthesizing high-fidelity, audio-driven video portraits with accurate emotional expressions. Our approach employs a variational autoencoder (VAE)-based audio-to-motion module to generate facial landmarks, which are concatenated with emotional embedding in a motion-to-emotion module to produce emotional landmarks. These landmarks drive a Neural Radiance Fields (NeRF)-based emotion-to-video module to render realistic emotional talking-head videos. Additionally, we propose a pose sampling method to generate natural idle-state (non-speaking) videos for silent audio inputs. Extensive experiments demonstrate that EmoGene outperforms previous methods in generating high-fidelity emotional talking-head videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。