构建真实且无法作弊的复杂多跳问答数据集,提升RAG系统评估可信度。
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
- 设计自动化流水线生成难以作弊的真实多跳问题。
- 在主流RAG模型上测试,作弊率降低81.0%。
- 适合评估和提升RAG系统在真实场景下的鲁棒性。
现实应用中,RAG系统常面临语料缺失或不完整的复杂查询。现有基准难以反映真实任务复杂性,常被通过非连贯推理(即无需真正多跳推理)作弊,或仅需简单事实回忆即可解决,限制了对现有RAG系统局限性的发现。为此,我们提出首个可自动、可控难度生成不可作弊、真实、无答案、多跳查询(CRUMQs)的流水线,适用于任意语料库和领域。我们在两个流行RAG数据集上构建CRUMQs,通过主流检索增强大模型的基准实验验证其有效性。结果表明,相比以往基准,CRUMQs对RAG系统极具挑战性,作弊率最高下降81.0%。该流水线为提升基准难度、推动更强大RAG系统的发展提供了简单有效的方法。
原文摘要 · Abstract (English)
Real-world use cases often present RAG systems with complex queries for which relevant information is missing from the corpus or is incomplete. In these settings, RAG systems must be able to reject unanswerable, out-of-scope queries and identify failures of retrieval and multi-hop reasoning. Despite this, existing RAG benchmarks rarely reflect realistic task complexity for multi-hop or out-of-scope questions, which often can be cheated via disconnected reasoning (i.e., solved without genuine multi-hop inference) or require only simple factual recall. This limits the ability for such benchmarks to uncover limitations of existing RAG systems. To address this gap, we present the first pipeline for automatic, difficulty-controlled creation of un$\underline{c}$heatable, $\underline{r}$ealistic, $\underline{u}$nanswerable, and $\underline{m}$ulti-hop $\underline{q}$uerie$\underline{s}$ (CRUMQs), adaptable to any corpus and domain. We use our pipeline to create CRUMQs over two popular RAG datasets and demonstrate its effectiveness via benchmark experiments on leading retrieval-augmented LLMs. Results show that compared to prior RAG benchmarks, CRUMQs are highly challenging for RAG systems and achieve up to 81.0\% reduction in cheatability scores. More broadly, our pipeline offers a simple way to enhance benchmark difficulty and drive development of more capable RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。