构建多模态心理理论评测基准,评估模型理解人类心理状态的能力。
MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
- 用短片叙事场景设计2300+多选题,覆盖7类心理理论任务。
- 视觉信息提升性能但模型难以有效融合多模态输入。
- 适合研究社会智能、多模态推理与具身认知的学者使用。
理解心理理论对构建能感知和解读人类行为的社会化多模态智能体至关重要。我们提出MoMentS(多模态心智状态),一个全面的基准,通过真实、叙事丰富的短片场景,评估多模态大语言模型(MLLMs)的心理理论能力。MoMentS包含超过2,300个多项选择题,涵盖七个不同的心理理论类别。该基准具备长视频上下文窗口和真实的社交互动,可深入揭示角色心智状态。我们评估了多个MLLMs,发现虽然视觉信息通常能提升表现,但模型仍难以有效整合多模态信息;处理音频对话的模型并未持续优于基于文本转录的输入。结果表明,亟需改进多模态融合,并指出了推动AI社会理解能力需克服的关键挑战。
原文摘要 · Abstract (English)
Understanding Theory of Mind is essential for building socially intelligent multimodal agents capable of perceiving and interpreting human behavior. We introduce MoMentS (Multimodal Mental States), a comprehensive benchmark designed to assess the ToM capabilities of multimodal large language models (LLMs) through realistic, narrative-rich scenarios presented in short films. MoMentS includes over 2,300 multiple-choice questions spanning seven distinct ToM categories. The benchmark features long video context windows and realistic social interactions that provide deeper insight into characters' mental states. We evaluate several MLLMs and find that although vision generally improves performance, models still struggle to integrate it effectively. For audio, models that process dialogues as audio do not consistently outperform transcript-based inputs. Our findings highlight the need to improve multimodal integration and point to open challenges that must be addressed to advance AI's social understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。