用AI自动评估RAG系统,准确率接近人工标注。
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
- 用大模型自动生成和匹配答案的关键词片段(nuggets)
- 自动评估与人工评估相关性达显著水平(21个主题,45次运行)
- 适合关注RAG评测方法改进的研究者和开发者
本报告展示了TREC 2024检索增强生成(RAG)赛道的部分初步结果。我们指出,RAG评估是信息获取及自然语言处理、人工智能领域持续进步的障碍,希望借此推动该领域的挑战解决。核心假设是:2003年为TREC问答赛道开发的关键词片段(nugget)评估法,可作为RAG系统评估的坚实基础。为此,我们提出“AutoNuggetizer”框架,利用大语言模型自动创建nuggets并将其匹配至系统回答。在TREC设置下,我们将完全自动流程与半人工的人类评估流程进行对比——由人工半手动创建nuggets,并手动分配给系统输出。基于21个主题、45次运行的初步结果,自动评估得分与人工评估得分呈现强相关性,表明全自动评估流程可用于指导未来RAG系统的迭代优化。
原文摘要 · Abstract (English)
This report provides an initial look at partial results from the TREC 2024 Retrieval-Augmented Generation (RAG) Track. We have identified RAG evaluation as a barrier to continued progress in information access (and more broadly, natural language processing and artificial intelligence), and it is our hope that we can contribute to tackling the many challenges in this space. The central hypothesis we explore in this work is that the nugget evaluation methodology, originally developed for the TREC Question Answering Track in 2003, provides a solid foundation for evaluating RAG systems. As such, our efforts have focused on "refactoring" this methodology, specifically applying large language models to both automatically create nuggets and to automatically assign nuggets to system answers. We call this the AutoNuggetizer framework. Within the TREC setup, we are able to calibrate our fully automatic process against a manual process whereby nuggets are created by human assessors semi-manually and then assigned manually to system answers. Based on initial results across 21 topics from 45 runs, we observe a strong correlation between scores derived from a fully automatic nugget evaluation and a (mostly) manual nugget evaluation by human assessors. This suggests that our fully automatic evaluation process can be used to guide future iterations of RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。