arXiv:2605.09703cs.CV2026-05中稿 · CVPR被引 1

构建真实场景心理状态理解数据集与多智能体框架,提升零样本推理能力

MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding

论文配图:MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding
图 1 · 摘自论文原文
  • 设计多模态真实学习场景数据集,标注行为、认知、情绪三类心理状态
  • 提出多智能体推理框架,零样本下心理状态预测准确率提升15.93点
  • 适合教育智能、人机交互领域研究者关注,推动真实场景心智理解

从自然行为中理解人类心理状态对真实世界智能系统至关重要。现有研究多聚焦孤立心理标签预测,缺乏复杂人际互动的结构化标注。为此,我们构建了精心设计的MOTOR-Bench基准,包含1,440个协作学习场景的多模态视频片段,涵盖自然类别不平衡、视觉噪声和领域特定语言等真实数据挑战。每条样本由教育专家基于自我调节学习理论标注。我们在MOTOR-Bench上评估多个先进多模态大模型与多智能体系统在零样本设置下的表现,结果表明其性能仍有限,说明现有方法在从可观察行为推断深层心理状态方面存在不足。为此,我们提出结构化多智能体推理框架MOTOR-MAS,通过协调机制分别推断显性行为、内部认知与心理情绪。实验显示,MOTOR-MAS在行为、认知、情绪三类标签的Macro-F1上比最优单模型提升15.93点,在内部认知预测上比通用多智能体基准提升10.2点。

原文摘要 · Abstract (English)

Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on predicting isolated mental state labels, lacking structured annotations of complex interpersonal interactions. To support structured analysis, we introduce MOTOR-Bench, a carefully-designed benchmark with a real-world dataset MOTOR-dataset, containing 1,440 multimodal video clips in collaborative learning scenarios, reflecting key real-world data challenges including natural class imbalance, visual noise, and domain-specific language. Each sample is labeled by educational experts based on self-regulated learning theory. We further evaluate several state-of-the-art multimodal large language models and multi-agent systems in a zero-shot setting on our MOTOR-Bench. However, their performance on this task remains limited, suggesting that existing methods still struggle with structured reasoning from observable behavior to deeper mental states. To address this challenge, we propose a reasoning multi-agent framework, named MOTOR-MAS. It coordinates multiple agents through a structured agent coordination mechanism to infer explicit behaviors, internal cognitions, and psychological emotions. Experimental results show that our MOTOR-MAS outperforms the best single-model benchmark by 15.93 points in Macro-F1 scores for the three labels of behavior, cognition, and emotion, and outperforms the general multi-agent benchmark by 10.2 points in internal cognition prediction.

心理状态理解多模态分析多智能体零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。