arXiv:2506.02387cs.AI2025-06被引 2

评测视觉语言模型在多智能体策略场景中的表现

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments

  • 构建10个视觉感知的多智能体环境,评估模型战略能力
  • 最佳模型仅达46.6%决策预测准确率和31.4%回报率
  • 揭示当前模型在推理与决策上的显著短板,适合研究多智能体系统者参考

视觉语言模型(VLMs)在交互式代理任务中取得进展,但现有基准仍局限于单智能体或纯文本环境。现实场景通常涉及多个智能体在丰富视觉与文本上下文中互动,对多模态感知和策略交互提出双重挑战。为此,我们提出视觉战略基准(VS-Bench),一个评估VLMs在多智能体环境中战略能力的多模态基准。VS-Bench包含十个视觉驱动的环境,涵盖合作、竞争与混合动机的互动。通过三个维度评估:感知(元素识别准确率)、策略推理(下一步动作预测准确率)与决策(归一化剧集回报)。对十五个领先VLMs的实验表明,尽管模型感知能力强,但在推理与决策上仍有显著差距,最优模型达到46.6%预测准确率与31.4%归一化回报。我们进一步分析影响性能的关键因素,开展人类实验并探究失败模式,以深化对VLMs战略能力的理解。通过标准化评估与揭示现有模型局限,我们期望VS-Bench成为未来战略多模态智能体研究的基础。代码与数据见https://vs-bench.github.io。

原文摘要 · Abstract (English)

Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain limited to single-agent or text-only environments. In contrast, real-world scenarios often involve multiple agents interacting within rich visual and textual contexts, posing challenges with both multimodal observations and strategic interactions. To bridge this gap, we introduce Visual Strategic Bench (VS-Bench), a multimodal benchmark that evaluates VLMs for strategic abilities in multi-agent environments. VS-Bench comprises ten vision-grounded environments that cover cooperative, competitive, and mixed-motive interactions. The performance of VLM agents is evaluated across three dimensions: perception measured by element recognition accuracy; strategic reasoning measured by next-action prediction accuracy; and decision-making measured by normalized episode return. Extensive experiments on fifteen leading VLMs show that, although current models exhibit strong perception abilities, there remains a significant gap to optimal performance in reasoning and decision-making, with the best-performing model attaining 46.6% prediction accuracy and 31.4% normalized return. We further analyze the key factors influencing performance, conduct human experiments, and examine failure modes to provide a deeper understanding of VLMs' strategic abilities. By standardizing the evaluation and highlighting the limitations of existing models, we envision VS-Bench as a foundation for future research on strategic multimodal agents. Code and data are available at https://vs-bench.github.io.

多智能体视觉语言模型战略推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。