构建人体语言心理理解基准,揭示当前AI在非言语互动中的认知短板。
Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
- 基于专家标注的身体语言库,构建细粒度视频数据集
- 现有AI在识别与解释非言语线索上表现远低于人类
- 适合研究具身智能、社会认知与人机交互的学者参考
通过非言语线索(NVCs)理解他人心理状态是生存与社会联结的基础。现有理论心智(ToM)评估多聚焦于错误信念任务和信息不对称推理,忽视了信念之外的心理状态及人类非言语交流的丰富性。本文提出Motion2Mind框架,用于评估机器在解读非言语线索方面的理论心智能力。依托专家精心整理的身体语言参考知识库,构建了一个精细标注的视频数据集,包含222类非言语线索与397种心理状态,并配有手动验证的心理学解释。评估显示,当前人工智能系统在识别与解释非言语线索方面存在显著性能差距,且相比人类标注者表现出过度解读现象。
原文摘要 · Abstract (English)
Our ability to interpret others' mental states through nonverbal cues (NVCs) is fundamental to our survival and social cohesion. While existing Theory of Mind (ToM) benchmarks have primarily focused on false-belief tasks and reasoning with asymmetric information, they overlook other mental states beyond belief and the rich tapestry of human nonverbal communication. We present Motion2Mind, a framework for evaluating the ToM capabilities of machines in interpreting NVCs. Leveraging an expert-curated body-language reference as a proxy knowledge base, we build Motion2Mind, a carefully curated video dataset with fine-grained nonverbal cue annotations paired with manually verified psychological interpretations. It encompasses 222 types of nonverbal cues and 397 mind states. Our evaluation reveals that current AI systems struggle significantly with NVC interpretation, exhibiting not only a substantial performance gap in Detection, as well as patterns of over-interpretation in Explanation compared to human annotators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。