构建多语言RAG评估基准,真实反映用户使用体验
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
- 基于原生语言问题与多模型生成响应,构建端到端评估框架
- 高一致标注结果验证了评估可靠性,揭示跨语言性能差异
- 适合开发和测试多语言自动评估器的开发者与研究者
检索增强生成(RAG)系统的自动评估依赖于忠实性与相关性等细粒度维度,需由专家人工标注。现有元评估基准多集中于英语或使用翻译数据,难以体现文化差异。本文提出多语言端到端元评估RAG基准MEMERAG,基于MIRACL数据集,采用原生语言问题,通过多种大语言模型生成回答,并由专家标注其忠实性与相关性。我们描述了标注流程并验证了高一致性。分析结果显示不同语言下生成模型表现存在差异。最后,我们将该数据集用于评估多语言自动评估器(如LLM-as-a-judge),证明其能可靠识别高级提示技术与模型带来的改进。数据集已开源。
原文摘要 · Abstract (English)
Automatic evaluation of retrieval augmented generation (RAG) systems relies on fine-grained dimensions like faithfulness and relevance, as judged by expert human annotators. Meta-evaluation benchmarks support the development of automatic evaluators that correlate well with human judgement. However, existing benchmarks predominantly focus on English or use translated data, which fails to capture cultural nuances. A native approach provides a better representation of the end user experience. In this work, we develop a Multilingual End-to-end Meta-Evaluation RAG benchmark (MEMERAG). Our benchmark builds on the popular MIRACL dataset, using native-language questions and generating responses with diverse large language models (LLMs), which are then assessed by expert annotators for faithfulness and relevance. We describe our annotation process and show that it achieves high inter-annotator agreement. We then analyse the performance of the answer-generating LLMs across languages as per the human evaluators. Finally we apply the dataset to our main use-case which is to benchmark multilingual automatic evaluators (LLM-as-a-judge). We show that our benchmark can reliably identify improvements offered by advanced prompting techniques and LLMs. Our dataset is available at https://github.com/amazon-science/MEMERAG
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。