基于真实视频生成自然协调的头眼运动,让视频更逼真
Data-driven Head Motion Generation through Natural Gaze-Head Coordination

- 用自动提取的真人头眼数据训练模型
- 生成的头动与视线自然匹配,且多样可信
- 适合做虚拟角色驱动、视频生成等场景
我们提出了首个基于数据驱动的方法,从大规模真实场景人脸视频中建模时序头眼协调关系。为获得可泛化的学习数据,我们设计了一套自动管道,利用现成的基于外观的眼动估计器提取自然且多样的眼动与头部运动。为捕捉头眼协调的概率相关性与时间动态,我们基于生成式条件变分自编码器构建模型,实现合理且多样的眼动引导头部运动生成。进一步,我们将框架应用于眼动控制的面部视频生成,实现了输入眼动驱动下自然真实的头部运动,这一方面此前未被重视。人类评估与定量对比表明,本方法有效且设计选择合理,评估者对本方法的偏好具有统计显著性。
原文摘要 · Abstract (English)
We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To obtain training data for generalizable learning, we propose an automatic pipeline that extracts natural yet diverse gaze and head motions with off-the-shelf appearance-based gaze estimators. To capture the probabilistic correlation and temporal dynamics of gaze-head coordination, we build our model on a generative conditional Variational Autoencoder for plausible yet diverse gaze-conditioned head motion generations. We further apply our framework to gaze-controlled facial video generation, where we enable video generation with natural and realistic head motion correlated to the input gaze - an aspect that has not been emphasized before. Human evaluation and quantitative comparisons demonstrate our method's effectiveness and validate our design choices, with evaluators showing statistically significant preference for our approach over baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。