arXiv:2502.15690cs.IRcs.AI2025-02被引 3

提出中文网页搜索代理框架与评测数据集,解决中文搜索模型评估不公问题。

Level-Navi Agent: A Framework and benchmark for Chinese Web Search Agents

  • 基于层级导航的无训练搜索代理,可跨多层网页理解复杂查询。
  • 构建标注数据集Web24,支持公平对比主流大模型性能。
  • 开源代码,助力中文AI搜索研究,适合中文NLP与智能代理方向研究者。

大型语言模型(LLMs)推动了人工智能网络搜索代理的发展,相比传统搜索引擎,其能更深入理解复杂查询并更好识别上下文。然而,中文网络搜索领域缺乏关注,导致开源模型能力未得到统一、公平评估。根本原因在于缺少三个关键要素:统一的代理框架、精确标注的数据集和合适的评估指标。为此,我们提出一种无需训练的通用网页搜索代理Level-Navi Agent,采用层级感知导航策略,能够逐层分析用户复杂问题并获取所需信息。同时,我们构建了高质量标注数据集Web24,并设计了合理的评估指标。在该标准下,对当前主流大模型进行了全面评估。为促进后续研究,源代码已公开于Github。

原文摘要 · Abstract (English)

Large language models (LLMs), adopted to understand human language, drive the development of artificial intelligence (AI) web search agents. Compared to traditional search engines, LLM-powered AI search agents are capable of understanding and responding to complex queries with greater depth, enabling more accurate operations and better context recognition. However, little attention and effort has been paid to the Chinese web search, which results in that the capabilities of open-source models have not been uniformly and fairly evaluated. The difficulty lies in lacking three aspects: an unified agent framework, an accurately labeled dataset, and a suitable evaluation metric. To address these issues, we propose a general-purpose and training-free web search agent by level-aware navigation, Level-Navi Agent, accompanied by a well-annotated dataset (Web24) and a suitable evaluation metric. Level-Navi Agent can think through complex user questions and conduct searches across various levels on the internet to gather information for questions. Meanwhile, we provide a comprehensive evaluation of state-of-the-art LLMs under fair settings. To further facilitate future research, source code is available at Github.

搜索代理中文NLP大模型评测网页导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。