自动构建鲁棒事实性评测集,揭示大模型在闭卷场景下的真实能力。
IRB: Automated Generation of Robust Factuality Benchmarks
- 用事实与算法双重支架自动生成评测数据
- 前沿大模型在闭卷设置下表现显著下降
- 推理模型更可靠,优化检索比扩容生成更划算
针对RAG系统静态评测基准易饱和且维护成本高的问题,我们提出IRB框架,实现事实性评测基准的自动化生成。IRB采用结构化生成流程,结合事实支架(factual scaffold)与算法支架(algorithmic scaffold)。利用该框架构建评测集,并对前沿大模型与检索器进行评估。结果表明,IRB在闭卷设置下对前沿大模型构成显著挑战。此外,评估显示推理型大模型更具可靠性,提升检索组件比扩大生成模型更能有效提高RAG系统的正确性。
原文摘要 · Abstract (English)
Static benchmarks for RAG systems often suffer from rapid saturation and require significant manual effort to maintain robustness. To address this, we present IRB, a framework for automatically generating benchmarks to evaluate the factuality of RAG systems. IRB employs a structured generation pipeline utilizing \textit{factual scaffold} and \textit{algorithmic scaffold}. We utilize IRB to construct a benchmark and evaluate frontier LLMs and retrievers. Our results demonstrate that IRB poses a significant challenge for frontier LLMs in the closed-book setting. Furthermore, our evaluation suggests that reasoning LLMs are more reliable, and that improving the retrieval component may yield more cost-effective gains in RAG system correctness than scaling the generator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。