发现大模型识表情时依赖牙齿等表面线索,可能产生偏差。
Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
- 用牙齿标注数据集测试视觉语言模型,分析其情绪判断依据。
- 模型在有牙时表现更好,但主要依赖眉毛等面部特征推断情绪。
- 揭示大模型存在快捷学习偏见,影响心理健康等敏感应用公平性。
基础模型(FMs)正快速改变情感计算领域,视觉语言模型(VLMs)已能在零样本场景下识别情绪。本文探究一个关键但未被充分关注的问题:这些模型依赖何种视觉线索推断情感?是心理上合理的依据,还是表层习得的捷径?我们在带有牙齿标注的AffectNet子集上对不同规模的VLM进行基准测试,发现模型性能随可见牙齿的存在而显著变化。通过对表现最佳的模型GPT-4o进行结构化内省,我们发现眉毛位置等面部属性驱动了其大部分情绪推理,其效价-唤醒度预测具有高度内部一致性。这些模式揭示了基础模型行为的涌现特性,但也暴露了潜在风险:捷径学习、偏见及公平性问题,尤其在心理健康与教育等敏感领域。
原文摘要 · Abstract (English)
Foundation Models (FMs) are rapidly transforming Affective Computing (AC), with Vision Language Models (VLMs) now capable of recognising emotions in zero shot settings. This paper probes a critical but underexplored question: what visual cues do these models rely on to infer affect, and are these cues psychologically grounded or superficially learnt? We benchmark varying scale VLMs on a teeth annotated subset of AffectNet dataset and find consistent performance shifts depending on the presence of visible teeth. Through structured introspection of, the best-performing model, i.e., GPT-4o, we show that facial attributes like eyebrow position drive much of its affective reasoning, revealing a high degree of internal consistency in its valence-arousal predictions. These patterns highlight the emergent nature of FMs behaviour, but also reveal risks: shortcut learning, bias, and fairness issues especially in sensitive domains like mental health and education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。