arXiv:2501.07572cs.CLcs.AI2025-01ACL被引 174

评测大模型网页导航能力,提升复杂信息获取效果

WebWalker: Benchmarking LLMs in Web Traversal

  • 设计多智能体框架模拟人类浏览网页,分步探索与评估
  • 在真实场景中实现垂直与水平信息整合,显著提升数据质量
  • 适合研究RAG与网页交互的学者及产品开发者

检索增强生成(RAG)在开放域问答任务中表现优异,但传统搜索引擎常返回浅层内容,限制大模型处理复杂、多层级信息的能力。为此,我们提出WebWalkerQA基准,用于评估大模型进行网页遍历的能力。该基准衡量模型系统性地访问网站子页面并提取高质量数据的性能。我们进一步提出WebWalker——一种基于‘探索-评判’范式的多智能体框架,模拟人类式网页导航。大量实验表明,WebWalkerQA具有挑战性,且结合WebWalker的RAG在真实场景中实现了横向与纵向的信息融合,有效提升了信息获取能力。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex, multi-layered information. To address it, we introduce WebWalkerQA, a benchmark designed to assess the ability of LLMs to perform web traversal. It evaluates the capacity of LLMs to traverse a website's subpages to extract high-quality data systematically. We propose WebWalker, which is a multi-agent framework that mimics human-like web navigation through an explore-critic paradigm. Extensive experimental results show that WebWalkerQA is challenging and demonstrates the effectiveness of RAG combined with WebWalker, through the horizontal and vertical integration in real-world scenarios.

大模型网页导航RAG多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。