arXiv:2604.08990cs.CV2026-04被引 1

让模型主动看脸,通过动态聚焦局部特征提升表情识别准确率。

ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning

  • 模型像侦探一样主动调用工具、聚焦关键面部区域进行观察。
  • 在FER任务中实现比传统方法更高的表情和动作单元预测精度。
  • 适合需要精细情绪分析的场景,如心理咨询与人机交互。

多模态大语言模型(MLLM)的发展为表情识别(FER)带来了新机遇,推动其从单纯标签预测转向基于推理的情感理解。然而,现有基于MLLM的FER方法仍采用被动模式:依赖预设人脸输入,仅对固定视觉证据进行单次推理,缺乏主动感知能力。为此,我们提出ActFER,一种将FER重构为动态视觉证据获取与多模态推理相结合的代理框架。具体而言,ActFER动态调用人脸检测与对齐工具,选择性放大信息丰富的局部区域,并通过视觉思维链(Chain-of-Thought)对动作单元(AUs)和情绪进行推理。为实现该行为,我们进一步设计了适配代理式FER的强化学习算法——效用校准的GRPO(UC-GRPO)。该算法利用基于动作单元的多层次可验证奖励增强监督,通过查询条件对比的效用估计实现样本感知的动态信用分配,以及情绪感知的EMA校准以降低噪声并捕捉情绪相关的检查倾向。该算法使ActFER能够学习何时进行局部检查以及如何基于获取的证据进行推理。全面实验表明,使用UC-GRPO训练的ActFER持续优于被动式基于MLLM的基准方法,并显著提升动作单元预测准确率。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have created new opportunities for facial expression recognition (FER), moving it beyond pure label prediction toward reasoning-based affect understanding. However, existing MLLM-based FER methods still follow a passive paradigm: they rely on externally prepared facial inputs and perform single-pass reasoning over fixed visual evidence, without the capability for active facial perception. To address this limitation, we propose ActFER, an agentic framework that reformulates FER as active visual evidence acquisition followed by multimodal reasoning. Specifically, ActFER dynamically invokes tools for face detection and alignment, selectively zooms into informative local regions, and reasons over facial Action Units (AUs) and emotions through a visual Chain-of-Thought. To realize such behavior, we further develop Utility-Calibrated GRPO (UC-GRPO), a reinforcement learning algorithm tailored to agentic FER. UC-GRPO uses AU-grounded multi-level verifiable rewards to densify supervision, query-conditional contrastive utility estimation to enable sample-aware dynamic credit assignment for local inspection, and emotion-aware EMA calibration to reduce noisy utility estimates while capturing emotion-wise inspection tendencies. This algorithm enables ActFER to learn both when local inspection is beneficial and how to reason over the acquired evidence. Comprehensive experiments show that ActFER trained with UC-GRPO consistently outperforms passive MLLM-based FER baselines and substantially improves AU prediction accuracy.

表情识别代理系统视觉推理强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。