评测大模型在模糊搜索中找网页的能力,发现现有系统表现普遍不佳。
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
- 设计新基准,用可调难度的模糊查询测试网页检索能力。
- 663个跨领域问题,多数模型准确率低于35%。
- 适合研究搜索代理、大模型推理与信息检索的学者。
大型语言模型已从简单聊天机器人演变为能自动化复杂现实任务的智能体,其浏览和推理实时网络内容的能力成为评估检索与认知技能的关键。现有基准如BrowseComp和xBench-DeepSearch侧重需要多跳推理的复杂查询,却忽视了模糊探索式搜索——用户提出模糊、多面向的问题,目标是找到最相关的网页而非单一事实答案。为此,我们提出Needle in the Web,一个专为评估现代搜索代理和基于LLM的系统在真实网络环境中应对模糊探索性查询而设计的新基准。该基准包含663个问题,覆盖七个不同领域,通过基于网页内容真实陈述的灵活方法生成可控难度的查询,确保问题质量与答案唯一性。我们在三个主流LLM和三个基于代理的搜索系统上进行测试,发现多数模型表现不佳:准确率普遍低于35%,且无一在所有领域或难度级别上持续领先。结果表明,当前搜索系统在语义模糊下的有效模糊检索仍面临重大挑战。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have evolved from simple chatbots into sophisticated agents capable of automating complex real-world tasks, where browsing and reasoning over live web content is key to assessing retrieval and cognitive skills. Existing benchmarks like BrowseComp and xBench-DeepSearch emphasize complex reasoning searches requiring multi-hop synthesis but neglect Fuzzy Exploratory Search, namely queries that are vague and multifaceted, where users seek the most relevant webpage rather than a single factual answer. To address this gap, we introduce Needle in the Web, a novel benchmark specifically designed to evaluate modern search agents and LLM-based systems on their ability to retrieve and reason over real-world web content in response to ambiguous, exploratory queries under varying levels of difficulty. Needle in the Web comprises 663 questions spanning seven distinct domains. To ensure high query quality and answer uniqueness, we employ a flexible methodology that reliably generates queries of controllable difficulty based on factual claims of web contents. We benchmark three leading LLMs and three agent-based search systems on Needle in the Web, finding that most models struggle: many achieve below 35% accuracy, and none consistently excel across domains or difficulty levels. These findings reveal that Needle in the Web presents a significant challenge for current search systems and highlights the open problem of effective fuzzy retrieval under semantic ambiguity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。