arXiv:2509.23698cs.CL2025-09EMNLP

构建人类中心场景决策评估基准,提升模型社会认知能力。

VIVA+: Human-Centered Situational Decision-Making

  • 基于真实场景设计三维度评测框架,聚焦情境理解与推理
  • 测试1317个场景6373道题,揭示主流模型在社会性决策中的短板
  • 适合关注具身智能、人机交互的开发者与研究者

多模态大语言模型在复杂人本环境中展现潜力,但评估其细腻的人类式推理与决策能力仍具挑战。本文提出VIVA+,一个基于认知理论的评测基准,涵盖1,317个真实世界情境和6,373道多选题,聚焦三大核心能力:基础情境理解、上下文驱动的动作解释与反思性推理。该框架系统评估模型在社会意义层面的感知、推理与行动能力。我们对最新商业与开源模型进行测评,揭示显著性能差异并指出关键挑战。通过针对性训练与多步推理策略,模型表现持续提升。深入分析揭示当前模型局限,为推进更具鲁棒性、上下文敏感且社交智能的多模态大模型提供可行路径。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) show promising results for embodied agents in operating meaningfully in complex, human-centered environments. Yet, evaluating their capacity for nuanced, human-like reasoning and decision-making remains challenging. In this work, we introduce VIVA+, a cognitively grounded benchmark for evaluating the reasoning and decision-making of MLLMs in human-centered situations. VIVA+ consists of 1,317 real-world situations paired with 6,373 multiple-choice questions, targeting three core abilities for decision-making: (1) Foundational Situation Comprehension, (2) Context-Driven Action Justification, and (3) Reflective Reasoning. Together, these dimensions provide a systematic framework for assessing a model's ability to perceive, reason, and act in socially meaningful ways. We evaluate the latest commercial and open-source models on VIVA+, where we reveal distinct performance patterns and highlight significant challenges. We further explore targeted training and multi-step reasoning strategies, which yield consistent performance improvements. Finally, our in-depth analysis highlights current model limitations and provides actionable insights for advancing MLLMs toward more robust, context-aware, and socially adept decision-making in real-world settings.

多模态模型决策评估人机交互具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。