arXiv:2511.00389cs.CV2025-11被引 7

用大模型统一表情识别,提升推理与可解释性

Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond

  • 将表情数据转为问答格式,用大模型直接推理
  • 构建23万条提示链数据,36万条强化学习数据
  • 新模型UniFER-7B在多个数据集上超越主流闭源模型

多模态大语言模型(MLLMs)已革新计算机视觉与情感计算领域。面部表情识别(FER)正从专用模型转向统一方法。一种有前景的路径是将传统FER数据集转化为视觉问答(VQA)格式,使通用大模型直接用于推理。然而,尽管顶尖MLLM在诸多任务中表现优异,其在FER任务上的性能仍缺乏系统评估。为此,我们构建了FERBench基准,涵盖4个常用FER数据集和20个先进MLLMs。结果表明,虽分类表现良好,但推理与可解释性仍存明显短板。为此,我们提出后训练策略:分别构建高质量大规模数据集UniFER-CoT-230K(用于冷启动初始化)和UniFER-RLVR-360K(用于可验证奖励的强化学习)。基于此,我们开发统一且可解释的FER基础模型UniFER-7B,其性能优于多个开源与闭源通用大模型(如Gemini-2.5-Pro和Qwen2.5-VL-72B)。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have revolutionized numerous research fields, including computer vision and affective computing. As a pivotal challenge in this interdisciplinary domain, facial expression recognition (FER) has evolved from separate, domain-specific models to more unified approaches. One promising avenue to unify FER tasks is converting conventional FER datasets into visual question-answering (VQA) formats, enabling the direct application of powerful generalist MLLMs for inference. However, despite the success of cutting-edge MLLMs in various tasks, their performance on FER tasks remains largely unexplored. To address this gap, we provide FERBench, a systematic benchmark that incorporates 20 state-of-the-art MLLMs across four widely used FER datasets. Our results reveal that, while MLLMs exhibit good classification performance, they still face significant limitations in reasoning and interpretability. To this end, we introduce post-training strategies aimed at enhancing the facial expression reasoning capabilities of MLLMs. Specifically, we curate two high-quality and large-scale datasets: UniFER-CoT-230K for cold-start initialization and UniFER-RLVR-360K for reinforcement learning with verifiable rewards (RLVR), respectively. Building upon them, we develop a unified and interpretable FER foundation model termed UniFER-7B, which outperforms many open-sourced and closed-source generalist MLLMs (e.g., Gemini-2.5-Pro and Qwen2.5-VL-72B).

表情识别大模型可解释性VQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。