测试视觉语言模型对模仿动作的理解能力,发现其表现远低于人类。
Can Vision Language Models Understand Mimed Actions?
- 构建基于动作捕捉的模仿动作评测基准MIME,包含86种动作。
- 模型在不同视角、角色和背景扰动下识别准确率显著下降。
- 适合研究具身交互、人机理解与多模态鲁棒性的学者参考。
非语言交流(NVC)在人类沟通中至关重要,但因其范围广且解释差异大而难以研究。而模仿表演——仅通过手势、表情和动作传递意图——是低解释变异性的一类典型行为。本文认为,掌握模仿动作是视觉语言模型理解更微妙非语言信号的前提。为此,提出基于视频问答的模仿动作评估基准MIME,包含86种模仿动作,利用动作捕捉数据构建,并引入角色、背景和视角扰动以评估模型鲁棒性。实验表明,无论是开源还是API调用的视觉语言模型,在MIME上的表现均显著低于人类,凸显当前模型在理解人体动作方面存在明显短板,亟需加强具身认知与动作理解的研究。
原文摘要 · Abstract (English)
Nonverbal communication (NVC) plays an integral role in human language, but studying NVC in general is challenging because of its broad scope and high variance in interpretation among individuals and cultures. However, mime -- the theatrical technique of suggesting intent using only gesture, expression, and movement -- is a subset of NVC that consists of explicit and embodied actions with much lower human interpretation variance. We argue that a solid understanding of mimed actions is a crucial prerequisite for vision-language models capable of interpreting and commanding more subtle aspects of NVC. Hence, we propose Mime Identification Multimodal Evaluation (MIME), a novel video-based question answering benchmark comprising of 86 mimed actions. Constructed with motion capture data, MIME consists of variations of each action with perturbations applied to the character, background, and viewpoint for evaluating recognition robustness. We find that both open-weight and API-based vision-language models perform significantly worse than humans on MIME, motivating the need for increased research for instilling more robust understanding of human gestures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。