评测大模型在真实网络中自主查信息的能力,发现现有方法有短板。
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
- 设计动态开放环境下的挑战性问题,模拟真实搜索场景
- 提出细粒度指标,评估信息获取的准确、有用与简洁程度
- 适合研究智能体搜索与检索增强生成的学者使用
检索增强生成(RAG)通过引入外部信息提升大语言模型(LLM)的表现。作为新兴范式,代理式RAG进一步将自主的LLM代理引入信息检索过程。然而,现有基准测试受限于静态检索环境和固定小规模语料库,且查询简单,无法激发代理行为。此外,其评估依赖预定义的文档黄金集,不适用于现实网络环境中开放动态的特点。为此,我们提出InfoDeepSeek,首个面向真实动态网络环境的代理式信息检索评测基准。我们建立系统化方法构建满足确定性、难度与多样性标准的挑战性问题,并开发首个专为动态代理信息寻求设计的评估框架,包含关于信息获取结果准确率、实用性与紧凑性的细粒度指标。在多种大模型、搜索引擎及问题类型上的广泛实验揭示了代理行为的细微差异,为未来研究提供可操作洞见。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by grounding responses with retrieved information. As an emerging paradigm, Agentic RAG further enhances this process by introducing autonomous LLM agents into the information seeking process. However, existing benchmarks fall short in evaluating such systems, as they are confined to a static retrieval environment with a fixed, limited corpus} and simple queries that fail to elicit agentic behavior. Moreover, their evaluation protocols assess information seeking effectiveness by pre-defined gold sets of documents, making them unsuitable for the open-ended and dynamic nature of real-world web environments. To bridge this gap, we present InfoDeepSeek, a new benchmark with challenging questions designed for assessing agentic information seeking in real-world, dynamic web environments. We propose a systematic methodology for constructing challenging queries satisfying the criteria of determinacy, difficulty, and diversity. Based on this, we develop the first evaluation framework tailored to dynamic agentic information seeking, including fine-grained metrics about the accuracy, utility, and compactness of information seeking outcomes. Through extensive experiments across LLMs, search engines, and question types, InfoDeepSeek reveals nuanced agent behaviors and offers actionable insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。