用情绪维度值让大模型理解脸表情,发现它更擅长描述而非分类。
Beyond Vision: How Large Language Models Interpret Facial Expressions from Valence-Arousal Values
- 用数值化的愉悦-唤醒值替代图像输入,测试大模型理解表情能力。
- 在描述任务中生成内容与人工标注高度一致,分类效果较差。
- 适合对情感生成、人机交互感兴趣的开发者和研究者。
大型语言模型主要依赖文本输入输出,但人类情绪通过言语和非言语线索(如面部表情)表达。尽管视觉-语言模型可分析图像中的面部表情,但其资源消耗大且可能过度依赖语言先验而非视觉理解。本研究探索大模型是否能从面部表情的数值化维度——愉悦度(Valence)和唤醒度(Arousal)值——推断情感含义。利用Facechannel从表情图像中提取VA值,并在两个任务中测试:(1) 将面部表情分类为基本情绪(IIMI数据集)和复杂情绪(Emotic数据集);(2) 生成表情的语义描述(Emotic数据集)。分类任务结果显示,大模型难以将VA值准确映射到离散情绪类别,尤其在超越基本极性(如快乐、悲伤)时表现不佳。但在语义描述任务中,模型生成的文本描述与人工标注高度一致,表明其在自由文本层面具有更强的情感推断能力。
原文摘要 · Abstract (English)
Large Language Models primarily operate through text-based inputs and outputs, yet human emotion is communicated through both verbal and non-verbal cues, including facial expressions. While Vision-Language Models analyze facial expressions from images, they are resource-intensive and may depend more on linguistic priors than visual understanding. To address this, this study investigates whether LLMs can infer affective meaning from dimensions of facial expressions-Valence and Arousal values, structured numerical representations, rather than using raw visual input. VA values were extracted using Facechannel from images of facial expressions and provided to LLMs in two tasks: (1) categorizing facial expressions into basic (on the IIMI dataset) and complex emotions (on the Emotic dataset) and (2) generating semantic descriptions of facial expressions (on the Emotic dataset). Results from the categorization task indicate that LLMs struggle to classify VA values into discrete emotion categories, particularly for emotions beyond basic polarities (e.g., happiness, sadness). However, in the semantic description task, LLMs produced textual descriptions that align closely with human-generated interpretations, demonstrating a stronger capacity for free text affective inference of facial expressions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。