arXiv:2409.02828cs.CVcs.MM2024-09被引 36

用大模型生成表情识别思维链,提升准确率与可解释性。

ExpLLM: Towards Chain of Thought for Facial Expression Recognition

  • 通过三步思维链分析表情:关键观察、情绪推理、最终结论。
  • 在 RAF-DB 与 AffectNet 上优于现有最先进方法,微表情识别更胜 GPT-4o。
  • 适合需要高可解释性的人机交互与情感计算研究者。

面部表情识别(FER)在多媒体领域具有重要意义,但分析表情成因对准确识别至关重要。现有基于面部动作单元(AUs)的方法通常仅提供名称与强度,缺乏对动作单元间互动关系及整体情绪的洞察。本文提出 ExpLLM,利用大语言模型生成精准的链式思维(CoT)用于表情识别。其思维链机制从三个层面构建:关键观察(描述 AU 名称、强度及关联情绪)、整体情绪推理(分析多 AU 互动,识别主导情绪及其关系)、最终结论(基于前序分析输出表情标签)。此外,我们设计了 Exp-CoT Engine,用于构建表达思维链并生成指令-描述数据以训练 ExpLLM。在 RAF-DB 与 AffectNet 数据集上的大量实验表明,ExpLLM 优于当前最先进方法,并在微表情识别中显著超越最新版 GPT-4o。

原文摘要 · Abstract (English)

Facial expression recognition (FER) is a critical task in multimedia with significant implications across various domains. However, analyzing the causes of facial expressions is essential for accurately recognizing them. Current approaches, such as those based on facial action units (AUs), typically provide AU names and intensities but lack insight into the interactions and relationships between AUs and the overall expression. In this paper, we propose a novel method called ExpLLM, which leverages large language models to generate an accurate chain of thought (CoT) for facial expression recognition. Specifically, we have designed the CoT mechanism from three key perspectives: key observations, overall emotional interpretation, and conclusion. The key observations describe the AU's name, intensity, and associated emotions. The overall emotional interpretation provides an analysis based on multiple AUs and their interactions, identifying the dominant emotions and their relationships. Finally, the conclusion presents the final expression label derived from the preceding analysis. Furthermore, we also introduce the Exp-CoT Engine, designed to construct this expression CoT and generate instruction-description data for training our ExpLLM. Extensive experiments on the RAF-DB and AffectNet datasets demonstrate that ExpLLM outperforms current state-of-the-art FER methods. ExpLLM also surpasses the latest GPT-4o in expression CoT generation, particularly in recognizing micro-expressions where GPT-4o frequently fails.

表情识别大模型思维链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。