arXiv:2508.13186cs.CLcs.AI2025-08被引 36

评测AI在图文视频混合网页中的智能浏览能力,发现顶级模型准确率仅24.25%。

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

  • 设计400个需结合图像视频的复杂问题,检验多模态信息提取能力。
  • 引入验证清单,可细粒度分析模型对视觉内容的依赖与推理路径。
  • 揭示当前顶尖模型在跨模态网页任务中仍表现不佳,适合评估多模态代理性能。

具备高级推理与工具使用能力的AI代理在深度网络搜索中表现出色,但现有基准(如BrowseComp)主要聚焦文本内容,忽视了多模态信息的普遍性。为此,我们提出MM-BrowseComp,一个包含400个精心设计的挑战性问题的新基准,用于评估多模态检索与推理能力。与以往工作不同,该基准引入视觉提示,要求从网页图片和视频中提取关键证据以完成任务,使纯文本方法失效。此外,我们为每个问题提供经验证的检查清单,支持对多模态依赖关系与推理路径的细粒度分析。对27个前沿模型的全面评估显示,即使是最先进的模型(如GPT-5-High带工具)也仅达到24.25%的准确率,凸显当前多模态网页浏览能力的不足,确立了本基准作为领域内严格新标准的地位。

原文摘要 · Abstract (English)

AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However, existing benchmarks such as BrowseComp primarily focus on textual content, overlooking the prevalence of multimodal content. To bridge this gap, we introduce MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities. Unlike prior work, MM-BrowseComp incorporates visual prompts and necessitates the extraction of key evidence from web images and videos to complete questions, rendering text-only approaches insufficient. Additionally, we provide a verified checklist for each question, enabling fine-grained analysis of multimodal dependencies and reasoning paths. Our comprehensive evaluation of 27 state-of-the-art models reveals that even leading models like GPT-5-High with tools achieve only 24.25\% accuracy, highlighting the suboptimal multimodal browsing capabilities, establishing MM-BrowseComp as a rigorous new standard for the field.

多模态网页代理评测基准视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。