用视觉语言模型生成身体反应情绪描述,识别遮脸图像中的情感。
Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Models
- 利用大视觉语言模型生成聚焦关键身体部位的情绪叙述。
- 在遮脸图像上识别情绪效果优于无微调的基线方法。
- 揭示模型对脸部区域的固有偏见,推动跨模态情感分析发展。
身体部位对情绪反应的体现包含丰富的情感体验信息。我们提出一个框架,利用先进的大型视觉语言模型(LVLMs)生成具身化情绪叙事(ELENA),即以显著身体部位为中心的多层次文本输出。通过注意力图观察发现,当前模型普遍存在对面部区域的持续偏好。尽管如此,该框架在未进行任何微调的情况下,仍能有效识别遮脸图像中的具身情绪,表现优于基线方法。ELENA为视觉与情感建模的跨模态研究开辟新路径,并增强情感感知环境下的建模能力。
原文摘要 · Abstract (English)
The embodiment of emotional reactions from body parts contains rich information about our affective experiences. We propose a framework that utilizes state-of-the-art large vision-language models (LVLMs) to generate Embodied LVLM Emotion Narratives (ELENA). These are well-defined, multi-layered text outputs, primarily comprising descriptions that focus on the salient body parts involved in emotional reactions. We also employ attention maps and observe that contemporary models exhibit a persistent bias towards the facial region. Despite this limitation, we observe that our employed framework can effectively recognize embodied emotions in face-masked images, outperforming baselines without any fine-tuning. ELENA opens a new trajectory for embodied emotion analysis across the modality of vision and enriches modeling in an affect-aware setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。