用大模型自动评估报告生成质量,效果接近人工。
Auto-ARGUE: LLM-Based Report Generation Evaluation
- 基于大模型实现ARGUE框架,自动评估报告生成质量。
- 在TREC 2024多个任务上与人工评分相关性良好。
- 提供可视化工具,支持细粒度分析报告缺陷。
引用支撑的报告生成是检索增强生成(RAG)系统的主要应用场景。尽管开源评估工具已覆盖多种RAG任务,但针对报告生成的工具仍匮乏。为此,我们提出Auto-ARGUE,一种基于大模型的、对近期提出的ARGUE框架的稳健实现。我们在TREC 2024 NeuCLIR赛道的报告生成预演任务,以及TREC 2024 RAG赛道的两个任务上对Auto-ARGUE进行了分析,结果显示其系统级评估结果与人工判断具有良好的相关性。此外,我们发布了ARGUE-Viz,一个用于可视化和细粒度分析Auto-ARGUE评分结果的Web应用。
原文摘要 · Abstract (English)
Generation of citation-backed reports is a primary use case for retrieval-augmented generation (RAG) systems. While open-source evaluation tools exist for various RAG tasks, tools designed for report generation are lacking. Accordingly, we introduce Auto-ARGUE, a robust LLM-based implementation of the recently proposed ARGUE framework for report generation evaluation. We present analysis of Auto-ARGUE on the report generation pilot task from the TREC 2024 NeuCLIR track and on two tasks from the TREC 2024 RAG track, showing good system-level correlations with human judgments. Additionally, we release ARGUE-Viz, a web app for visualization and fine-grained analysis of Auto-ARGUE judgments and scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。