用真人偏好优化面部表情,让数字人更自然生动。
SuperFace: Preference-Aligned Facial Expression Estimation Beyond Pseudo Supervision

- 以真人对表情的偏好反馈替代伪标签,优化面部动作预测
- 在真实人脸动画中显著提升表达真实感,优于传统方法
- 适合需要高感知真实性的数字人、影视动画领域
准确的面部表情估计对实现逼真的数字人动画至关重要,ARKit的blendshape系数提供了可解释的语义化动画控制表示。然而,高质量的ARKit系数预测受限于缺乏可靠的真实标注。现有方法通常依赖Live Link Face等捕捉软件生成伪标签,但这些标签存在噪声激活、系数幅度偏差以及面部动作缺失或不准等问题。因此,基于监督学习的模型往往复制不完美的伪标签,而非追求感知上的表达保真度。本文提出SuperFace,一种偏好驱动的框架,将ARKit面部表情估计从伪标签模仿转向人类感知对齐优化。不把软件估算系数当作固定真值,而是仅用作初始化,再通过人类对渲染表情的偏好反馈进一步改进系数预测。通过对感知判断对齐而非数值伪标签,SuperFace实现了更具视觉真实性和表现力的面部动画。实验表明,相比Live Link Face的监督方式,SuperFace在表达保真度上取得显著提升,验证了偏好驱动优化在语义面部动作预测中的有效性。
原文摘要 · Abstract (English)
Accurate facial estimation is crucial for realistic digital human animation, and ARKit blendshape coefficients offer an interpretable representation by mapping facial motions to semantic animation controls. However, learning high-quality ARKit coefficient prediction remains limited by the absence of reliable ground-truth supervision. Existing methods typically rely on capture software such as Live Link Face to provide pseudo labels, which may contain noisy activations, biased coefficient magnitudes, and missing or inaccurate facial actions. Consequently, models trained with supervised learning tend to reproduce imperfect pseudo labels rather than optimize for perceptual expression fidelity. In this paper, we propose SuperFace, a preference-driven framework that moves ARKit facial expression estimation from pseudo-label imitation toward human-aligned perceptual optimization. Instead of treating software-estimated coefficients as fixed ground truth, SuperFace uses them only as an initialization and further improves coefficient prediction through human preference feedback on rendered facial expressions. By aligning the model with perceptual judgments rather than numerical pseudo labels, SuperFace enables more visually faithful and expressive facial animation. Experiments show that SuperFace improves expression fidelity over Live Link Face supervision, demonstrating the effectiveness of preference-driven optimization for semantic facial action prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。