提出语义级表情表示,实现野外人脸表情的鲁棒捕捉与跨身份迁移。
SEREP: Semantic Facial Expression Representation for Robust In-the-Wild Capture and Retargeting
- 在语义层面解耦表情与身份,提升表达表征的泛化性。
- 利用低质量合成数据训练,在单目图像上准确预测表情。
- 构建多身份表情基准数据集MultiREX,推动该领域评估标准化。
单目野外人脸表演捕捉因拍摄条件差异、脸型多样及表情复杂而极具挑战。现有方法多依赖线性3D可变形模型,在顶点位移层面独立表示表情与身份。本文提出语义表情表示(SEREP),在语义层面实现表情与身份的解耦。首先基于高质量3D非配对表情数据学习表达表示;随后通过一种新颖的半监督方案,利用低质量合成数据训练模型从单目图像中预测表情。此外,本文构建了MultiREX基准数据集,填补表情捕捉任务缺乏评测资源的空白。实验表明,SEREP优于现有最先进方法,在捕捉复杂表情及跨身份迁移方面表现卓越。
原文摘要 · Abstract (English)
Monocular facial performance capture in-the-wild is challenging due to varied capture conditions, face shapes, and expressions. Most current methods rely on linear 3D Morphable Models, which represent facial expressions independently of identity at the vertex displacement level. We propose SEREP (Semantic Expression Representation), a model that disentangles expression from identity at the semantic level. We start by learning an expression representation from high-quality 3D data of unpaired facial expressions. Then, we train a model to predict expression from monocular images relying on a novel semi-supervised scheme using low quality synthetic data. In addition, we introduce MultiREX, a benchmark addressing the lack of evaluation resources for the expression capture task. Our experiments show that SEREP outperforms state-of-the-art methods, capturing challenging expressions and transferring them to new identities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。