评测视觉主导的智能体多步深度推理能力,发现顶级模型成功率不足50%。
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
- 构建包含828个真实场景任务的多模态基准,覆盖视觉、视频、文本等
- 顶尖模型在多步任务中成功率低于50%,暴露推理与工具使用瓶颈
- 适合研究视觉智能体、多步推理和具身智能的学者参考
深度推理对解决复杂任务至关重要,尤其在需要序列化、多模态理解的视觉主导场景中。然而,现有基准通常仅评估合成的单轮查询,视觉模态有限,且缺乏对多步骤推理质量的评估框架。为此,我们提出Agent-X,一个大规模基准,用于评估视觉主导智能体在真实多模态环境中的多步深度推理能力。Agent-X包含828个智能体任务,涵盖图像、多图对比、视频及说明文本等真实视觉上下文,覆盖六大智能体环境:通用视觉推理、网页浏览、安防监控、自动驾驶、体育分析和数学推理。任务要求智能体在不同场景中结合工具使用与显式分步决策。我们还设计了细粒度的步骤级评估框架,评估每一步的正确性、逻辑连贯性及工具使用有效性。结果表明,即使是最优模型(如GPT、Gemini、Qwen系列)在多步视觉任务中的全链路成功率为50%以下。该发现揭示了当前大模型在视觉推理与工具调用上的关键瓶颈,并指明未来研究方向。数据与代码已公开于https://github.com/mbzuai-oryx/Agent-X。
原文摘要 · Abstract (English)
Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully synthetic, single-turn queries, limited visual modalities, and lack a framework to assess reasoning quality over multiple steps as required in real-world settings. To address this, we introduce Agent-X, a large-scale benchmark for evaluating vision-centric agents multi-step and deep reasoning capabilities in real-world, multimodal settings. Agent- X features 828 agentic tasks with authentic visual contexts, including images, multi-image comparisons, videos, and instructional text. These tasks span six major agentic environments: general visual reasoning, web browsing, security and surveillance, autonomous driving, sports, and math reasoning. Our benchmark requires agents to integrate tool use with explicit, stepwise decision-making in these diverse settings. In addition, we propose a fine-grained, step-level evaluation framework that assesses the correctness and logical coherence of each reasoning step and the effectiveness of tool usage throughout the task. Our results reveal that even the best-performing models, including GPT, Gemini, and Qwen families, struggle to solve multi-step vision tasks, achieving less than 50% full-chain success. These findings highlight key bottlenecks in current LMM reasoning and tool-use capabilities and identify future research directions in vision-centric agentic reasoning models. Our data and code are publicly available at https://github.com/mbzuai-oryx/Agent-X
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。