arXiv:2506.19079cs.CVcs.AI2025-06被引 3

发现大模型识表情时依赖牙齿等表面线索,可能产生偏差。

Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition

  • 用牙齿标注数据集测试视觉语言模型,分析其情绪判断依据。
  • 模型在有牙时表现更好,但主要依赖眉毛等面部特征推断情绪。
  • 揭示大模型存在快捷学习偏见,影响心理健康等敏感应用公平性。

基础模型(FMs)正快速改变情感计算领域,视觉语言模型(VLMs)已能在零样本场景下识别情绪。本文探究一个关键但未被充分关注的问题:这些模型依赖何种视觉线索推断情感?是心理上合理的依据,还是表层习得的捷径?我们在带有牙齿标注的AffectNet子集上对不同规模的VLM进行基准测试,发现模型性能随可见牙齿的存在而显著变化。通过对表现最佳的模型GPT-4o进行结构化内省,我们发现眉毛位置等面部属性驱动了其大部分情绪推理,其效价-唤醒度预测具有高度内部一致性。这些模式揭示了基础模型行为的涌现特性,但也暴露了潜在风险:捷径学习、偏见及公平性问题,尤其在心理健康与教育等敏感领域。

原文摘要 · Abstract (English)

Foundation Models (FMs) are rapidly transforming Affective Computing (AC), with Vision Language Models (VLMs) now capable of recognising emotions in zero shot settings. This paper probes a critical but underexplored question: what visual cues do these models rely on to infer affect, and are these cues psychologically grounded or superficially learnt? We benchmark varying scale VLMs on a teeth annotated subset of AffectNet dataset and find consistent performance shifts depending on the presence of visible teeth. Through structured introspection of, the best-performing model, i.e., GPT-4o, we show that facial attributes like eyebrow position drive much of its affective reasoning, revealing a high degree of internal consistency in its valence-arousal predictions. These patterns highlight the emergent nature of FMs behaviour, but also reveal risks: shortcut learning, bias, and fairness issues especially in sensitive domains like mental health and education.

情绪识别大模型偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。