arXiv:2607.26465cs.AI2026-07中稿 · EMNLP

评测多模态模型在故事中的动机推理能力,发现现有模型难以持续理解行为动因变化。

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

论文配图:MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
图 1 · 摘自论文原文
  • 基于马斯洛与雷思理论构建故事驱动的多模态动机推理评估框架。
  • 所有测试模型在连续上下文中均无法保持一致的动机推理,表现显著不足。
  • 适合关注社会智能、动态推理与人类行为理解的研究者使用。

多模态大语言模型因其在社交智能方面的潜力而受到广泛关注,但其在序列动机推理方面的能力仍缺乏深入研究。现有评估大多聚焦静态文本或孤立视觉片段,无法反映现实行为动因的累积特性。为此,我们提出 MultivationBench,一个面向故事驱动视觉叙事的多模态动机推理基准。该基准基于马斯洛需求层次和雷思基本欲望理论,要求模型整合累积的多模态上下文以推断不断演变的动机。结果表明,MultivationBench 构成重大挑战:所有测试模型在连续上下文中均难以维持一致的动机推理,暴露出静态识别能力与实现类人社交理解所需动态推理之间的关键脱节。

原文摘要 · Abstract (English)

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.

多模态动机推理序列理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。