通过语义表达参数生成带情绪的口型视频,控制更精准。
EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters
- 用音频标签预测情绪表达参数,提升跨情绪相关性。
- 借助预训练超平面优化面部动作,生成质量更高。
- 适合需要精细情绪控制的虚拟人/交互系统应用。
从音频输入生成带有特定情绪的说话头视频是人机交互中的重要且复杂挑战。由于情绪是高度抽象且边界模糊的概念,需解耦表达参数以生成富有情感的表现。本文提出 EmoHead,通过语义表达参数合成说话头视频。为预测任意音频输入的表达参数,引入一个可由情绪标签指定的音频-表达模块,旨在增强不同情绪间音频输入的相关性。此外,利用预训练超平面沿垂直方向探测,进一步优化面部运动。最终,经过精炼的表达参数用于正则化神经辐射场,促进情绪一致的说话头视频生成。实验表明,语义表达参数能显著提升重建质量和可控性。
原文摘要 · Abstract (English)
Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled expression parameters to generate emotionally expressive talking head videos. In this work, we present EmoHead to synthesize talking head videos via semantic expression parameters. To predict expression parameter for arbitrary audio input, we apply an audio-expression module that can be specified by an emotion tag. This module aims to enhance correlation from audio input across various emotions. Furthermore, we leverage pre-trained hyperplane to refine facial movements by probing along the vertical direction. Finally, the refined expression parameters regularize neural radiance fields and facilitate the emotion-consistent generation of talking head videos. Experimental results demonstrate that semantic expression parameters lead to better reconstruction quality and controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。