新基准评估长文档检索生成,更准衡量大模型用信息能力
Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall
- 构建含280题的长文本检索评测集,每题配5篇平均2444词文档
- 提出关键点召回率指标,量化模型对检索信息的利用程度
- 适合研究长文本生成、信息整合或评估大模型知识运用的研究者
检索增强生成(RAG)是解决大语言模型(LLM)知识固定性的有效方法。然而,现有评估基准存在两大缺陷:一是缺乏能反映真实检索文档特征的数据集,难以有效评估模型处理长上下文的能力;二是缺乏对长文本生成中有效利用检索信息的全面评估方法。为此,本文提出Long²RAG基准与关键点召回(KPR)指标。Long²RAG包含280个跨10个领域、8类问题的题目,每题关联5篇平均2,444词的文档。KPR通过检测生成回答中是否包含从检索文档提取的关键点,提供对模型信息利用能力的更精细评估。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) is a promising approach to address the limitations of fixed knowledge in large language models (LLMs). However, current benchmarks for evaluating RAG systems suffer from two key deficiencies: (1) they fail to adequately measure LLMs' capability in handling long-context retrieval due to a lack of datasets that reflect the characteristics of retrieved documents, and (2) they lack a comprehensive evaluation method for assessing LLMs' ability to generate long-form responses that effectively exploits retrieved information. To address these shortcomings, we introduce the Long$^2$RAG benchmark and the Key Point Recall (KPR) metric. Long$^2$RAG comprises 280 questions spanning 10 domains and across 8 question categories, each associated with 5 retrieved documents with an average length of 2,444 words. KPR evaluates the extent to which LLMs incorporate key points extracted from the retrieved documents into their generated responses, providing a more nuanced assessment of their ability to exploit retrieved information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。