用哑剧视频训练AI理解非语言社交,突破纯语言模型局限。
MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models
- 构建哑剧视频问答数据集MimeQA,聚焦非语言社交推理。
- 现有视频大模型在哑剧任务上准确率仅20-30%,人类达86%。
- 适合研究具身智能、人机交互与多模态理解的学者参考。
随着人工智能日益融入日常生活,具备社会智能、能自然理解并互动于人类日常的AI变得愈发重要。然而,当前AI社会推理研究普遍依赖语言或以语言为主的方法,导致系统在言语交流方面进步显著,却在非语言社交理解上表现不佳。为此,本文利用富含非语言社交互动的哑剧视频作为新数据源。哑剧通过肢体动作和姿态表达,为解读非语言社交沟通带来独特挑战与机遇。我们构建了名为MimeQA的新数据集,从YouTube获取约8小时视频片段,并设计涵盖806个精心标注与验证的问答对,用于评测非语言社交推理能力。基于MimeQA,我们评估了当前最先进的视频大语言模型(VideoLLMs),发现其准确率普遍仅为20-30%,而人类达到86%。分析显示,这些模型常无法正确锚定想象中的物体,过度依赖文本提示,忽视细微的非语言互动。本工作旨在推动未来具备真正社会智能、能理解非语言人类互动的AI模型发展。
原文摘要 · Abstract (English)
As AI becomes more closely integrated with peoples' daily activities, socially intelligent AI that can understand and interact seamlessly with humans in daily lives is increasingly important. However, current works in AI social reasoning all rely on language-only or language-dominant approaches to benchmark and training models, resulting in systems that are improving in verbal communication but struggle with nonverbal social understanding. To address this limitation, we tap into a novel data source rich in nonverbal social interactions -- mime videos. Mimes refer to the art of expression through gesture and movement without spoken words, which presents unique challenges and opportunities in interpreting nonverbal social communication. We contribute a new dataset called MimeQA, obtained by sourcing ~8 hours of videos clips from YouTube and developing a comprehensive video question-answering benchmark comprising 806 carefully annotated and verified question-answer pairs, designed to probe nonverbal social reasoning capabilities. Using MimeQA, we evaluate state-of-the-art video large language models (VideoLLMs) and find that they achieve low accuracy, generally ranging from 20-30%, while humans score 86%. Our analysis reveals that VideoLLMs often fail to ground imagined objects and over-rely on the text prompt while ignoring subtle nonverbal interactions. We hope to inspire future work in AI models that embody true social intelligence capable of interpreting non-verbal human interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。