用语音生成多样自然的3D眼神动作,解决眼神与语音关联弱的问题。
TalkingEyes: Pluralistic Speech-Driven 3D Eye Gaze Animation
- 构建14小时高质量语音-眼神数据集,融合头动与面部动作
- 双潜空间联合生成头动与眼神,降低语音与非语言动作关联建模难度
- 可同步生成眼神、眨眼、头部及面部动作,适合虚拟人与动画应用
尽管近期语音驱动3D面部动画取得显著进展,但一个关键的面部组件——眼神方向却常被忽视。这主要源于语音与眼神间关联较弱,且缺乏音频-眼神数据。本文提出一种新颖的数据驱动方法,可生成与语音协调的多样化3D眼神动作。首先,我们构建了一个约14小时的语音-网格序列数据集,包含高质量的眼神、头部和面部运动,通过轻量级眼神拟合与人脸重建从现有音视频数据集中提取。随后,设计了一种新的语音到动作转换框架,将头动与眼神动作为一体生成,但在两个独立潜空间中建模。该设计基于生理学知识:眼球转动范围小于头部。通过将语音嵌入映射至两个潜空间,有效缓解了语音与非语言动作间弱相关性带来的建模难题。最终,TalkingEyes结合语音驱动3D面部动作生成器,可从语音协同合成眼神、眨眼、头部与面部动作。大量定量与定性评估表明,该方法在生成多样且自然的3D眼神动作方面表现优异。
原文摘要 · Abstract (English)
Although significant progress has been made in the field of speech-driven 3D facial animation recently, the speech-driven animation of an indispensable facial component, eye gaze, has been overlooked by recent research. This is primarily due to the weak correlation between speech and eye gaze, as well as the scarcity of audio-gaze data, making it very challenging to generate 3D eye gaze motion from speech alone. In this paper, we propose a novel data-driven method which can generate diverse 3D eye gaze motions in harmony with the speech. To achieve this, we firstly construct an audio-gaze dataset that contains about 14 hours of audio-mesh sequences featuring high-quality eye gaze motion, head motion and facial motion simultaneously. The motion data is acquired by performing lightweight eye gaze fitting and face reconstruction on videos from existing audio-visual datasets. We then tailor a novel speech-to-motion translation framework in which the head motions and eye gaze motions are jointly generated from speech but are modeled in two separate latent spaces. This design stems from the physiological knowledge that the rotation range of eyeballs is less than that of head. Through mapping the speech embedding into the two latent spaces, the difficulty in modeling the weak correlation between speech and non-verbal motion is thus attenuated. Finally, our TalkingEyes, integrated with a speech-driven 3D facial motion generator, can synthesize eye gaze motion, eye blinks, head motion and facial motion collectively from speech. Extensive quantitative and qualitative evaluations demonstrate the superiority of the proposed method in generating diverse and natural 3D eye gaze motions from speech. The project page of this paper is: https://lkjkjoiuiu.github.io/TalkingEyes_Home/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。