arXiv:2502.13465cs.IRcs.AI2025-02NeurIPS被引 9

构建多领域分层评估基准,测试RAG系统在复杂查询下的适应能力。

HawkBench: Investigating Resilience of RAG Methods on Stratified Information-Seeking Tasks

  • 按信息获取行为分层任务类型,覆盖事实与推理类查询
  • 包含1600个标注样本,跨领域均衡分布以减少数据偏差
  • 揭示RAG需动态决策与全局理解才能提升泛化性能

在真实世界的信息检索场景中,用户需求动态多样,要求RAG系统具备自适应的鲁棒性。为全面评估现有RAG方法的韧性,我们提出HawkBench——一个由人工标注、跨多领域的基准测试集,用于系统评估RAG在不同任务类型下的表现。通过基于信息获取行为对任务进行分层,HawkBench能有效衡量RAG系统对多样化用户需求的适应能力。相比现有基准(主要聚焦于事实类查询且依赖不同知识库),HawkBench具备三大优势:(1) 系统化的任务分层设计,涵盖事实型与推理型查询;(2) 所有任务类型均整合多领域语料库,降低语料偏见;(3) 严格标注确保评估质量。该基准包含1600个高质量测试样本,均匀分布在各领域和任务类型中。利用此基准,我们评估了代表性RAG方法在答案质量和响应延迟方面的表现。结果表明,亟需融合决策机制、查询理解与全局知识认知的动态策略,以提升RAG的泛化能力。我们认为HawkBench将成为推动RAG系统韧性及通用信息检索能力发展的关键基准。

原文摘要 · Abstract (English)

In real-world information-seeking scenarios, users have dynamic and diverse needs, requiring RAG systems to demonstrate adaptable resilience. To comprehensively evaluate the resilience of current RAG methods, we introduce HawkBench, a human-labeled, multi-domain benchmark designed to rigorously assess RAG performance across categorized task types. By stratifying tasks based on information-seeking behaviors, HawkBench provides a systematic evaluation of how well RAG systems adapt to diverse user needs. Unlike existing benchmarks, which focus primarily on specific task types (mostly factoid queries) and rely on varying knowledge bases, HawkBench offers: (1) systematic task stratification to cover a broad range of query types, including both factoid and rationale queries, (2) integration of multi-domain corpora across all task types to mitigate corpus bias, and (3) rigorous annotation for high-quality evaluation. HawkBench includes 1,600 high-quality test samples, evenly distributed across domains and task types. Using this benchmark, we evaluate representative RAG methods, analyzing their performance in terms of answer quality and response latency. Our findings highlight the need for dynamic task strategies that integrate decision-making, query interpretation, and global knowledge understanding to improve RAG generalizability. We believe HawkBench serves as a pivotal benchmark for advancing the resilience of RAG methods and their ability to achieve general-purpose information seeking.

RAG评估基准信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。