arXiv:2511.10254cs.CV2025-11被引 3

让情绪识别与推理对齐,提升表情分析的准确与可信度。

Facial-R1: Aligning Reasoning and Recognition for Facial Emotion Analysis

  • 分三阶段对齐识别与推理过程,减少错误解释。
  • 在8个基准上达领先效果,测试集准确率超现有方法。
  • 适合需要可解释情绪分析的医疗、人机交互场景。

面部情绪分析(FEA)通过引入可解释的细粒度推理,扩展了传统情绪识别任务,整合了情绪识别、面部动作单元(AU)识别和基于AU的情绪推理三个子任务,联合建模情感状态。尽管近期方法利用视觉语言模型(VLMs)取得良好成果,但仍存在两大关键问题:(1)生成性幻觉,即模型因缺乏特定情绪知识而产生看似合理但不准确的解释;(2)情绪推理与识别之间的错位,源于观测面部特征与最终标签间连接断裂。我们提出Facial-R1,一种三阶段对齐框架,以最小监督解决上述问题。首先通过指令微调建立基础情绪推理能力;其次引入强化学习,以情绪和AU标签作为奖励信号,显式对齐推理过程与预测情绪;第三,设计数据合成管道,迭代利用前序阶段扩充训练数据,实现模型的可扩展自提升。基于该框架,我们构建了FEA-20K基准数据集,包含17,737条训练样本和1,688条测试样本,具备细粒度标注。在八个标准基准上的实验表明,Facial-R1在FEA任务中达到最先进性能,具备强泛化能力和鲁棒可解释性。

原文摘要 · Abstract (English)

Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrates three subtasks: emotion recognition, facial Action Unit (AU) recognition, and AU-based emotion reasoning to model affective states jointly. While recent approaches leverage Vision-Language Models (VLMs) and achieve promising results, they face two critical limitations: (1) hallucinated reasoning, where VLMs generate plausible but inaccurate explanations due to insufficient emotion-specific knowledge; and (2) misalignment between emotion reasoning and recognition, caused by fragmented connections between observed facial features and final labels. We propose Facial-R1, a three-stage alignment framework that effectively addresses both challenges with minimal supervision. First, we employ instruction fine-tuning to establish basic emotional reasoning capability. Second, we introduce reinforcement training guided by emotion and AU labels as reward signals, which explicitly aligns the generated reasoning process with the predicted emotion. Third, we design a data synthesis pipeline that iteratively leverages the prior stages to expand the training dataset, enabling scalable self-improvement of the model. Built upon this framework, we introduce FEA-20K, a benchmark dataset comprising 17,737 training and 1,688 test samples with fine-grained emotion analysis annotations. Extensive experiments across eight standard benchmarks demonstrate that Facial-R1 achieves state-of-the-art performance in FEA, with strong generalization and robust interpretability.

情绪分析可解释性多任务学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。