arXiv:2504.20119cs.IRcs.AI2025-04中稿 · presentation at th…综述被引 16

用大模型评估大模型生成的检索系统,靠谱吗?

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets

  • 用大模型自动构建数据集并评估RAG各组件表现
  • 调研63篇论文,发现自动化评估可行但需改进
  • 适合想落地RAG的企业参考评估规范

近年来,检索增强生成(RAG)系统取得显著进展。由于其包含索引、检索和生成等多个组件及大量参数,系统性评估与质量提升面临巨大挑战。以往研究指出,评估RAG系统对记录进展、比较配置、识别特定领域有效方法至关重要。本研究系统回顾了63篇学术论文,全面梳理当前RAG评估方法,重点关注四大方向:数据集、检索器、索引与数据库、生成组件。我们发现,利用具备生成与评估能力的大模型,可实现对RAG各组件的自动化评估。同时,亟需更多实践研究,为公司提供实施与评估RAG系统的明确指导。通过整合关键组件的评估方法,并强调领域专用数据集的创建与适配,我们推动了系统化评估方法的发展与评估严谨性的提升。此外,通过对基于大模型的自动化方法与人工判断的互动分析,我们参与了自动化与人工输入平衡的讨论,厘清二者在实现稳健可靠评估中的贡献、局限与挑战。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) has advanced significantly in recent years. The complexity of RAG systems, which involve multiple components-such as indexing, retrieval, and generation-along with numerous other parameters, poses substantial challenges for systematic evaluation and quality enhancement. Previous research highlights that evaluating RAG systems is essential for documenting advancements, comparing configurations, and identifying effective approaches for domain-specific applications. This study systematically reviews 63 academic articles to provide a comprehensive overview of state-of-the-art RAG evaluation methodologies, focusing on four key areas: datasets, retrievers, indexing and databases, and the generator component. We observe the feasibility of an automated evaluation approach for each component of a RAG system, leveraging an LLM capable of both generating evaluation datasets and conducting evaluations. In addition, we found that further practical research is essential to provide companies with clear guidance on the do's and don'ts of implementing and evaluating RAG systems. By synthesizing evaluation approaches for key RAG components and emphasizing the creation and adaptation of domain-specific datasets for benchmarking, we contribute to the advancement of systematic evaluation methods and the improvement of evaluation rigor for RAG systems. Furthermore, by examining the interplay between automated approaches leveraging LLMs and human judgment, we contribute to the ongoing discourse on balancing automation and human input, clarifying their respective contributions, limitations, and challenges in achieving robust and reliable evaluations.

RAG评估大模型评测自动化评估领域数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。