构建首个图文交错搜索基准,测试多轮视觉证据融合能力。
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

- 设计三层次图文交替搜索任务,支持主动寻图与多分支比对
- 最先进模型准确率不足50%,暴露出视觉推理整合难题
- 提供标准化工具链,适合评估多模态智能体的搜索能力
现有图文搜索评测仅涵盖单次视觉浏览,未将视觉证据融入多轮交互式搜索流程。本文提出InterLV-Search,首个面向语言-视觉协同搜索的基准,支持文本与视觉证据在搜索过程中反复迭代使用。包含2,061个样本,分三个层级:主动视觉证据获取、受控离线交错搜索、开放网络交错搜索;新增多分支对比任务,模拟真实场景中对多个实体的并行证据检索。前两级通过自动化流程构建,第三级采用机器主导、人工监督的开放式网页采集。同时提供InterLV-Agent工具,实现标准工具调用、轨迹记录与评估。在自研及开源多模态代理上测试显示,当前系统整体准确率均低于50%,凸显视觉证据搜寻、搜索控制与多模态融合等核心挑战。数据与代码已开源至https://github.com/hbhalpha/InterLV-Search-Bench。
原文摘要 · Abstract (English)
Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoint rather than part of an interleaved search trajectory. We introduce \textbf{InterLV-Search}, a benchmark for Interleaved Language-Vision Agentic Search, in which textual and visual evidence is repeatedly used to condition later search. It contains 2,061 examples across three levels: active visual evidence seeking, controlled offline interleaved multimodal search, and open-web interleaved multimodal search. Beyond existing benchmarks, it also includes multimodal multi-branch samples that involve comparison between multiple entities during the evidence search. We construct Level 1 and Level 2 with automated pipelines and Level 3 with a machine-led, human-supervised open-web pipeline. We further provide InterLV-Agent for standardized tool use, trajectory logging, and evaluation. Experiments on proprietary and open-source multimodal agents show that current systems remain far from solving interleaved multimodal search, with the best model below 50% overall accuracy, highlighting challenges in visual evidence seeking, search control, and multimodal evidence integration. We release the benchmark data and evaluation code at https://github.com/hbhalpha/InterLV-Search-Bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。