arXiv:2604.03016cs.AI2026-04被引 3

新基准评估多模态智能体真实推理过程,揭示现有模型在复杂任务中表现严重不足。

Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?

  • 设计可验证过程的评测框架,支持工具调用与双轴轨迹追踪
  • 模型在最难题目上准确率仅23.0%,远低于整体56.3%水平
  • 适合研究智能体能力、评估工具使用效率的学者参考

多模态大语言模型正从被动观察者转向主动智能体,通过视觉扩展(调用视觉工具)和知识扩展(开放式网络搜索)解决问题。然而现有评估存在缺陷:缺乏灵活工具集成、分立测试视觉与搜索工具、仅以最终答案评分,无法验证工具是否被调用、正确使用或高效执行。为此,我们提出Agentic-MME,一个面向多模态智能体能力的过程验证基准。该基准包含418个真实世界任务,覆盖6个领域与3个难度等级,每项任务平均需超过10人时的手动标注,共生成超2000个步骤级检查点。每个任务配备统一评估框架,支持沙箱代码与API调用,并提供人类参考轨迹,标注双轴(S轴与V轴)步骤级节点。为实现真正过程级验证,我们审计细粒度中间状态,而非仅看最终答案,并通过相对于人类轨迹的过思考度量效率。实验表明,最佳模型Gemini3-pro总体准确率为56.3%,但在三级任务上骤降至23.0%,凸显真实多模态智能体问题求解的挑战性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall short: they lack flexible tool integration, test visual and search tools separately, and evaluate primarily by final answers. Consequently, they cannot verify if tools were actually invoked, applied correctly, or used efficiently. To address this, we introduce Agentic-MME, a process-verified benchmark for Multimodal Agentic Capabilities. It contains 418 real-world tasks across 6 domains and 3 difficulty levels to evaluate capability synergy, featuring over 2,000 stepwise checkpoints that average 10+ person-hours of manual annotation per task. Each task includes a unified evaluation framework supporting sandboxed code and APIs, alongside a human reference trajectory annotated with stepwise checkpoints along dual-axis: S-axis and V-axis. To enable true process-level verification, we audit fine-grained intermediate states rather than just final answers, and quantify efficiency via an overthinking metric relative to human trajectories. Experimental results show the best model, Gemini3-pro, achieves 56.3% overall accuracy, which falls significantly to 23.0% on Level-3 tasks, underscoring the difficulty of real-world multimodal agentic problem solving.

多模态智能体过程验证评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。