MERRIN benchmark测试搜索智能体在嘈杂网络中跨模态推理能力,发现现有模型表现远低于人类。
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

- 设计多模态噪声环境下的检索与推理任务,模拟真实复杂搜索场景
- 平均准确率仅22.3%,最强模型也未超40.1%,暴露现有系统严重不足
- 适合研究多模态搜索、智能体推理或评估真实网络环境下AI性能的学者
针对搜索查询的模糊性和多跳特性,以及真实网络结果的多模态、异构性与冲突性,我们提出MERRIN(多模态证据检索与推理在噪声网络环境中的基准),一个由人工标注的搜索增强型智能体评估基准。MERRIN衡量智能体识别相关模态、检索多模态证据,并在噪声网页源上进行多跳推理的能力。其区别于以往工作体现在三方面:(1)使用无显式模态提示的自然语言查询;(2)引入视频、音频等未充分探索的模态;(3)要求从复杂、常含噪声或矛盾信息的网页中检索证据。我们在三种搜索设置下评估了十种模型(包括GPT-5.4-mini、Gemini 3/3.1 Flash/Pro等闭源模型及Qwen3-4B/30B/235B等开源模型)的性能。结果显示,所有智能体平均准确率为22.3%,最佳模型仅达40.1%。尽管更强模型如Gemini Deep Research表现更好,但提升有限,因过度探索导致被冲突或部分相关的内容干扰。相比人类,这些智能体消耗更多资源却准确率更低,主要由于源选择效率低和对文本模态过度依赖。研究凸显了在噪声环境中实现鲁棒跨模态搜索与推理的必要性,使MERRIN成为评估此类能力的重要基准。
原文摘要 · Abstract (English)
Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web results, we introduce MERRIN (Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments), a human-annotated benchmark for evaluating search-augmented agents. MERRIN measures AI agents' ability to identify relevant modalities, retrieve multimodal evidence, and perform multi-hop reasoning over noisy web sources. It differs from prior work in three important aspects: (1) using natural language queries without explicit modality cues, (2) incorporating underexplored modalities such as video and audio, and (3) requiring the retrieval of complex, often noisy or conflicting multimodal evidence during web search. We evaluate diverse search agents powered by ten models, including strong closed-source models (e.g., GPT-5.4-mini, Gemini 3/3.1 Flash/Pro) and open-weight models (Qwen3-4B/30B/235B), across three search settings (no search, native search, and agentic search). Our results show that MERRIN is highly challenging: the average accuracy across all agents is 22.3%, with the best-performing agent reaching only 40.1%. We further observe that while stronger agents like Gemini Deep Research achieve higher performance, gains are modest due to over-exploration; they take more steps and use more tools, but are often distracted by conflicting or partially relevant web content, leading to incorrect answers. Compared to humans, these agents consume more resources yet achieve lower accuracy, largely due to inefficient source selection and an overreliance on text modalities. These findings highlight the need for search agents capable of robust search and reasoning across diverse modalities in noisy web environments, making MERRIN a valuable testbed for evaluating such capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。