评测多模态智能体在真实复杂视觉场景中的长程工具协作能力
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
- 构建跨7大类25子领域的真实视觉任务,融合多模态工具链
- 顶尖模型仅达27.3%准确率,难题需超25步工具调用
- 适合研究通用智能体、长程推理与多模态系统者参考
现实世界中的多模态智能体需基于视觉证据完成多步骤任务流程。例如,通过图像匹配电路图并验证修复方案,或解读交通图规划行程并满足路线约束。然而现有基准多聚焦单轮视觉推理或特定工具技能,难以体现实际应用所需的现实性、视觉细微差别及长周期工具使用。我们提出AgentVista,一个面向通用多模态智能体的基准,涵盖7大类别25个子领域,结合高细节度的真实视觉场景与自然混合工具使用。任务要求跨模态的长程工具交互,包括网页搜索、图像搜索、页面导航及基于代码的图像处理与通用编程操作。对当前领先模型的全面评估揭示了显著差距:即使最佳模型Gemini-3-Pro(带工具)整体准确率也仅27.3%,部分难题需超过25次工具调用。我们期望AgentVista能推动更强大、可靠的多模态智能体在真实复杂场景中的发展。
原文摘要 · Abstract (English)
Real-world multimodal agents solve multi-step workflows grounded in visual evidence. For example, an agent can troubleshoot a device by linking a wiring photo to a schematic and validating the fix with online documentation, or plan a trip by interpreting a transit map and checking schedules under routing constraints. However, existing multimodal benchmarks mainly evaluate single-turn visual reasoning or specific tool skills, and they do not fully capture the realism, visual subtlety, and long-horizon tool use that practical agents require. We introduce AgentVista, a benchmark for generalist multimodal agents that spans 25 sub-domains across 7 categories, pairing realistic and detail-rich visual scenarios with natural hybrid tool use. Tasks require long-horizon tool interactions across modalities, including web search, image search, page navigation, and code-based operations for both image processing and general programming. Comprehensive evaluation of state-of-the-art models exposes significant gaps in their ability to carry out long-horizon multimodal tool use. Even the best model in our evaluation, Gemini-3-Pro with tools, achieves only 27.3% overall accuracy, and hard instances can require more than 25 tool-calling turns. We expect AgentVista to accelerate the development of more capable and reliable multimodal agents for realistic and ultra-challenging problem solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。