arXiv:2505.13419cs.CV2025-05被引 14

构建情感协同推理的多模态大模型,提升面部表情分析精度与泛化能力。

FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning

  • 设计新型指令数据集,对齐表情与动作单元并建立因果关系。
  • 在多个数据集上实现零样本泛化,性能优于现有方法。
  • 适合需要高可解释性情感分析的研究者和开发者。

面部情绪分析(FEA)在视觉情感计算中至关重要,旨在基于面部数据推断人的心理状态。科学上,面部表情由面部肌肉协同运动产生,可分解为特定动作单元(AUs),提供细致的情绪信息。然而传统方法普遍存在可解释性差、泛化能力弱及推理能力不足的问题。近期,多模态大语言模型(MLLMs)在多种视觉任务中表现优异,但在FEA任务中仍受限于缺乏专用数据集及其难以捕捉表情与动作单元间复杂关系。为此,我们构建了一个新的高质量面部表情指令数据集,精准对齐表情与动作单元描述,并建立其因果推理关系,同时创建新基准FEABench。进一步提出FEALLM,一种专为捕捉更细粒度面部信息而设计的新型MLLM架构。该模型在FEABench上表现优异,并在多个公开数据集(RAF-DB、AffectNet、BP4D、DISFA)上通过零样本评估展现出强大泛化能力,验证了其在FEA任务中的鲁棒性与有效性。数据集与代码将开源于https://github.com/953206211/FEALLM。

原文摘要 · Abstract (English)

Facial Emotion Analysis (FEA) plays a crucial role in visual affective computing, aiming to infer a person's emotional state based on facial data. Scientifically, facial expressions (FEs) result from the coordinated movement of facial muscles, which can be decomposed into specific action units (AUs) that provide detailed emotional insights. However, traditional methods often struggle with limited interpretability, constrained generalization and reasoning abilities. Recently, Multimodal Large Language Models (MLLMs) have shown exceptional performance in various visual tasks, while they still face significant challenges in FEA due to the lack of specialized datasets and their inability to capture the intricate relationships between FEs and AUs. To address these issues, we introduce a novel FEA Instruction Dataset that provides accurate and aligned FE and AU descriptions and establishes causal reasoning relationships between them, followed by constructing a new benchmark, FEABench. Moreover, we propose FEALLM, a novel MLLM architecture designed to capture more detailed facial information, enhancing its capability in FEA tasks. Our model demonstrates strong performance on FEABench and impressive generalization capability through zero-shot evaluation on various datasets, including RAF-DB, AffectNet, BP4D, and DISFA, showcasing its robustness and effectiveness in FEA tasks. The dataset and code will be available at https://github.com/953206211/FEALLM.

情感分析多模态大模型面部识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。