构建强制多模态推理的搜索基准,提升视觉线索追踪能力。
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- 通过迭代图文检索与交叉验证,强制模型依赖视觉细节
- 311个任务中最强模型端到端准确率36.0%,引入标记模块后最高提3.9分
- 适合研究多模态代理、视觉推理与真实场景搜索的开发者
现有多模态浏览基准常因可仅用文本启发式解题而缺乏真实多模态推理。我们提出MMSearch-Plus,一个包含311个任务的基准,通过要求在检索噪声下进行细粒度视觉线索提取与传播,强制实现多模态理解。数据集设计引导回答需从空间线索和时间轨迹推断出图像外事实,如事件、日期与地点。我们还提供无模型依赖的代理框架与一组标记(SoM)模块,支持标记、裁剪子区域及定向检索。SoM实现溯源感知的缩放与检索,提升多步推理鲁棒性。我们在该框架下评估了闭源与开源多模态大模型。最强系统端到端准确率为36.0%,引入SoM在多个场景中均带来稳定增益,最高达+3.9点。失败分析显示,定位相关网页与区分视觉相似事件仍是主要错误来源。这些结果凸显现实多模态搜索挑战,并确立MMSearch-Plus作为推进代理型多模态大模型的严格基准。
原文摘要 · Abstract (English)
Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verification. We introduce MMSearch-Plus, a 311-task benchmark that enforces multimodal understanding by requiring extraction and propagation of fine-grained visual cues through iterative image-text retrieval and cross-validation under retrieval noise. Our curation procedure seeds questions whose answers require extrapolating from spatial cues and temporal traces to out-of-image facts such as events, dates, and venues. Beyond the dataset, we provide a model-agnostic agent framework with standard browsing tools and a set-of-mark (SoM) module, which lets the agent place marks, crop subregions, and launch targeted image/text searches. SoM enables provenance-aware zoom-and-retrieve and improves robustness in multi-step reasoning. We evaluated closed- and open-source MLLMs in this framework. The strongest system achieves an end-to-end accuracy of 36.0%, and integrating SoM produces consistent gains in multiple settings, with improvements up to +3.9 points. From failure analysis, we observe recurring errors in locating relevant webpages and distinguishing between visually similar events. These results underscore the challenges of real-world multimodal search and establish MMSearch-Plus as a rigorous benchmark for advancing agentic MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。