arXiv:2603.14827cs.CV2026-03被引 2

让人脸动作估计结果可解释,用语言模型提升准确性与泛化能力

SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space

  • 将系数预测转为语义推理,通过两阶段知识蒸馏构建可解释空间
  • 在真实与卡通人脸上均实现更高精度和感知一致性,跨身份泛化强
  • 适合需要可控、可解释表情生成的虚拟角色与人机交互场景

从单张图像进行人脸动作估计通常被建模为在紧凑表达空间中预测或拟合参数,但此类方法缺乏明确的语义可解释性。而许多实际应用如虚拟角色控制和人机交互,需要对应于有意义肌肉运动的可解释面部动作。本文提出SemanticFace框架,在可解释的ARKit blendshape空间中实现面部动作估计,将系数预测重构为结构化语义推理。该方法采用两阶段语义蒸馏范式:首先从真实ARKit系数中提取结构化语义监督信号,再将其知识蒸馏至多模态大语言模型,从而从图像中预测可解释的面部动作系数。大量实验表明,语言对齐的语义监督显著提升了系数准确性和感知一致性,同时具备强大的跨身份泛化能力与对大规模域偏移(包括卡通人脸)的鲁棒性。

原文摘要 · Abstract (English)

Facial action estimation from a single image is often formulated as predicting or fitting parameters in compact expression spaces, which lack explicit semantic interpretability. However, many practical applications, such as avatar control and human-computer interaction, require interpretable facial actions that correspond to meaningful muscle movements. In this work, we propose SemanticFace, a framework for facial action estimation in the interpretable ARKit blendshape space that reformulates coefficient prediction as structured semantic reasoning. SemanticFace adopts a two-stage semantic distillation paradigm: it first derives structured semantic supervision from ground-truth ARKit coefficients and then distills this knowledge into a multimodal large language model to predict interpretable facial action coefficients from images. Extensive experiments demonstrate that language-aligned semantic supervision improves both coefficient accuracy and perceptual consistency, while enabling strong cross-identity generalization and robustness to large domain shifts, including cartoon faces.

人脸动作估计可解释性语言模型跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。