实测四种RAG评估指标在真实业务问答中的有效性
Evaluating RAG Metrics in Applied Contexts: An Experiment, Its Findings and Its Limitations
- 用真实业务数据构建问答集,测试四类RAG评估工具
- 发现多数指标与人工评分相关性较弱,部分与召回率不一致
- 适合关注RAG评估可靠性的研究人员和工程团队
本文报告了一项实证研究,评估多个RAG评估指标在实际应用中的相关性。实验基于从商业数据中由人工标注者构建的问答数据集,对RAG系统生成的回答及检索片段,采用来自四个库(Ragas、DeepEval、RAGChecker、Opik)的评估指标进行评分,并与两名评估员的打分以及标准指标(如召回率)进行对比。通过相关性分析,揭示了现有指标在实际场景中的局限性。最后,论文指出本方法的不足之处,与文献中常用方法进行比较,并提出未来研究方向。本文为原法语会议EvalLLM(Brabant, 2026)论文的英文翻译。
原文摘要 · Abstract (English)
This paper reports an empirical study evaluating the relevance of several RAG metrics. The experiment is based on a question-answering dataset created by human annotators from business data. The generated responses and retrieved spans of a RAG system are scored using evaluation metrics from four libraries (Ragas, DeepEval, RAGChecker, Opik). These metrics are compared to scores given by two evaluators, as well as to standard metrics such as recall. An analysis of correlations is conducted. Finally, we highlight certain limitations of our methodology, compare it to those used in the literature, and suggest some avenues for future research. This paper is an English translation of a paper originally published in the French-speaking workshop EvalLLM (Brabant, 2026).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。