arXiv:2504.14891cs.CL2025-04综述被引 51

系统梳理大模型时代RAG评估方法与数据集,助力提升生成准确性与可靠性。

Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey

  • 构建融合检索与生成的评估框架,覆盖性能、事实性、安全性和效率
  • 整理并分类10余类RAG专用数据集与评估工具,涵盖主流研究场景
  • 分析高影响力论文的评估实践,为开发者提供可复用的方法论

近年来,检索增强生成(RAG)通过将大语言模型(LLM)与外部信息检索结合,推动自然语言处理发展,实现跨场景的精准、实时且可验证的文本生成。然而,由于其混合架构及对动态知识源的依赖,在大模型时代评估RAG系统面临独特挑战。本文系统综述了传统与新兴的RAG评估方法与框架,涵盖系统性能、事实准确性、安全性与计算效率等方面。我们整理并分类了十余类RAG专用数据集与评估工具,并对高影响力研究中的评估实践进行元分析。据我们所知,这是目前最全面的RAG评估综述,连接传统与大模型驱动的评估方法,为推进RAG技术发展提供关键资源。

原文摘要 · Abstract (English)

Recent advancements in Retrieval-Augmented Generation (RAG) have revolutionized natural language processing by integrating Large Language Models (LLMs) with external information retrieval, enabling accurate, up-to-date, and verifiable text generation across diverse applications. However, evaluating RAG systems presents unique challenges due to their hybrid architecture that combines retrieval and generation components, as well as their dependence on dynamic knowledge sources in the LLM era. In response, this paper provides a comprehensive survey of RAG evaluation methods and frameworks, systematically reviewing traditional and emerging evaluation approaches, for system performance, factual accuracy, safety, and computational efficiency in the LLM era. We also compile and categorize the RAG-specific datasets and evaluation frameworks, conducting a meta-analysis of evaluation practices in high-impact RAG research. To the best of our knowledge, this work represents the most comprehensive survey for RAG evaluation, bridging traditional and LLM-driven methods, and serves as a critical resource for advancing RAG development.

RAG评估大模型文本生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。