arXiv:2511.21245cs.CV2025-11

让3D人脸重建更懂表情,直接监督提升情感识别准确率

FIELDS: Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision

  • 用2D/3D联合监督直接学习表情代码,避免传统方法的强度放大陷阱
  • 在内部和外部数据集上均提升情感识别效果,几何还原精度仍保持领先
  • 适合需要精准表达分析的面部情绪理解与交互应用

单目3D人脸重建从单张图像中估计3D可变形模型(3DMM)表示,生成具备几何感知的表情编码,对表情分析与情感理解具有价值。尽管进展显著,多数方法采用图像级自监督训练,评估主要关注几何保真度,未必最大化所学表情表示的情感效用,且在情感监督粗略耦合时可能引发强度放大捷径。本文提出FIELDS(基于直接监督的精准表情推断人脸重建),一个任务驱动框架,在几何合理性约束下学习FLAME表情代码以优化面部表情识别(FER)。通过混合2D/3D监督,FIELDS在域内与域外评估中均提升情感预测性能,同时在保留的和跨域3D基准上保持竞争力的几何保真度。

原文摘要 · Abstract (English)

Monocular 3D face reconstruction estimates a 3D morphable model (3DMM) representation from a single image, providing geometry-aware expression codes that are useful for facial expression analysis and affect understanding. Despite strong progress, most pipelines are trained with image-level self-supervision and evaluated primarily by geometric fidelity, which does not necessarily maximize the affective utility of the learned expression representation and may encourage intensity-amplifying shortcuts when affect supervision is naively coupled. We propose FIELDS (Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision), a task-driven framework that learns FLAME expression codes for facial expression recognition (FER) under a geometric plausibility constraint. Using hybrid 2D/3D supervision, FIELDS improves affect prediction in both in-domain and external evaluations while maintaining competitive geometric fidelity on held-out and out-of-domain 3D benchmarks.

3D人脸重建表情识别直接监督几何保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。