arXiv:2504.13128cs.IRcs.AI2025-04NeurIPS被引 21

构建真实技术文档检索评估基准,发现现有模型表现远低于理想水平。

FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents

  • 自动收集代码与技术文档,生成社区真实问题及答案
  • 五项新数据集挑战现有模型,显著低于理想检索效果
  • 揭示重排序器无效场景,验证优质上下文对RAG生成的关键作用

我们提出FreshStack,一种全自动构建信息检索(IR)评估基准的综合框架,融合具有挑战性的问答对。该框架包含三步:(1) 自动从代码和专业技术文档中收集语料库;(2) 基于社区提问与回答生成知识片段(nugget);(3) 采用多种检索技术融合与混合架构,在片段级别进行支持度判定。我们基于快速发展的新兴、近期及小众主题构建了五个数据集,确保任务足够具有挑战性。在这些数据集上,现有检索模型未经调整直接应用时,在全部五个主题上的表现均显著落后于理想方法(oracle),表明仍有巨大提升空间。此外,我们发现重排序器在其中两个主题上未能提升首阶段检索准确率,且优质上下文能显著帮助大语言模型生成高质量RAG答案。我们希望FreshStack能推动未来构建更真实、可扩展、无污染的IR与RAG评估基准的研究。

原文摘要 · Abstract (English)

We introduce FreshStack, a holistic framework for automatically building information retrieval (IR) evaluation benchmarks by incorporating challenging questions and answers. FreshStack conducts the following steps: (1) automatic corpus collection from code and technical documentation, (2) nugget generation from community-asked questions and answers, and (3) nugget-level support, retrieving documents using a fusion of retrieval techniques and hybrid architectures. We use FreshStack to build five datasets on fast-growing, recent, and niche topics to ensure the tasks are sufficiently challenging. On FreshStack, existing retrieval models, when applied out-of-the-box, significantly underperform oracle approaches on all five topics, denoting plenty of headroom to improve IR quality. In addition, we identify cases where rerankers do not improve first-stage retrieval accuracy (two out of five topics) and oracle context helps an LLM generator generate a high-quality RAG answer. We hope FreshStack will facilitate future work toward constructing realistic, scalable, and uncontaminated IR and RAG evaluation benchmarks.

信息检索RAG评估基准技术文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。