arXiv:2410.13206cs.CL2024-10ACL被引 4

构建人体语言问答数据集,测试视频大模型对肢体情绪的识别能力

BQA: Body Language Question Answering Dataset for Video Large Language Models

  • 构建包含26种情绪标签的肢体语言问答数据集BQA
  • 多款视频大模型在该数据集上表现不佳,平均准确率不足50%
  • 发现模型对年龄和种族存在显著偏见,易误判非裔与老年群体

人类交流中大量依赖面部表情、眼神接触和肢体语言等非语言线索。与语言或手语不同,这些非语言沟通缺乏正式规则,需要基于常识进行复杂推理。让当前视频大语言模型(VideoLLMs)准确理解肢体语言是一项关键挑战,因为人类无意识动作容易导致模型误判意图。为此,我们提出了一个名为BQA的肢体语言问答数据集,用于验证模型能否从短时肢体语言视频中正确识别情绪,涵盖26种情绪类别。我们在BQA上评估了多种VideoLLMs,发现理解肢体语言仍具挑战性;对错误答案的分析显示,部分模型在面对不同年龄和族裔的视频人物时表现出明显偏见。

原文摘要 · Abstract (English)

A large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language. Unlike language or sign language, such nonverbal communication lacks formal rules, requiring complex reasoning based on commonsense understanding. Enabling current Video Large Language Models (VideoLLMs) to accurately interpret body language is a crucial challenge, as human unconscious actions can easily cause the model to misinterpret their intent. To address this, we propose a dataset, BQA, a body language question answering dataset, to validate whether the model can correctly interpret emotions from short clips of body language comprising 26 emotion labels of videos of body language. We evaluated various VideoLLMs on BQA and revealed that understanding body language is challenging, and our analyses of the wrong answers by VideoLLMs show that certain VideoLLMs made significantly biased answers depending on the age group and ethnicity of the individuals in the video. The dataset is available.

视频理解情绪识别偏见检测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。