构建首个分级难度的RAG评估数据集,助力系统化评测问答模型性能。
LiveRAG: A diverse Q&A dataset with varying difficulty level for RAG evaluation
- 基于竞赛数据合成895个问答对,涵盖不同难度层级。
- 引入真实答案与支持性依据,实现精准效果评估。
- 用项目反应理论量化题目难度,区分模型能力差异。
随着检索增强生成(RAG)在生成式AI中的日益重要,系统化评估其有效性成为迫切需求。本文提出LiveRAG基准,一个公开可获取的合成问答数据集,包含895个合成问题与答案,源自SIGIR'2025 LiveRAG挑战赛。该数据集补充了比赛期间未公开的信息,包括真实答案及其支持性陈述,用于评估模型输出质量。每道题还附带由项目反应理论(Item Response Theory)模型计算出的难度与辨别度评分,揭示问题多样性及对系统能力的区分能力。分析表明,该基准能有效评估不同RAG系统的性能表现,推动社区在系统性评测和鲁棒性问答系统方面的研究进展。
原文摘要 · Abstract (English)
With Retrieval Augmented Generation (RAG) becoming more and more prominent in generative AI solutions, there is an emerging need for systematically evaluating their effectiveness. We introduce the LiveRAG benchmark, a publicly available dataset of 895 synthetic questions and answers designed to support systematic evaluation of RAG-based Q&A systems. This synthetic benchmark is derived from the one used during the SIGIR'2025 LiveRAG Challenge, where competitors were evaluated under strict time constraints. It is augmented with information that was not made available to competitors during the Challenge, such as the ground-truth answers, together with their associated supporting claims which were used for evaluating competitors' answers. In addition, each question is associated with estimated difficulty and discriminability scores, derived from applying an Item Response Theory model to competitors' responses. Our analysis highlights the benchmark's questions diversity, the wide range of their difficulty levels, and their usefulness in differentiating between system capabilities. The LiveRAG benchmark will hopefully help the community advance RAG research, conduct systematic evaluation, and develop more robust Q&A systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。