评估尽职调查中RAG系统的可靠性,发现并量化其幻觉等缺陷。
Towards a rigorous evaluation of RAG systems: the challenge of due diligence
- 结合人工与LLM评分,构建可靠评估流程
- 在投资尽调场景中检测出系统幻觉与引用失败
- 为工业级RAG评估提供可复用的协议与数据集
生成式AI的兴起推动了医疗、金融等高风险领域的发展。检索增强生成(RAG)架构通过融合语言模型(LLMs)与搜索引擎,能够基于文档库生成回答,具有显著潜力。然而,在关键应用场景中,RAG系统的可靠性仍存疑,幻觉等问题持续存在。本研究评估了一个用于投资基金尽职调查的RAG系统。我们提出一种结合人工标注与LLM-Judge标注的稳健评估协议,用于识别系统故障,如幻觉、偏离主题、引用失败和回避回答。受预测驱动推理(PPI)方法启发,实现具有统计保证的精确性能测量。我们提供一个全面的数据集以支持进一步分析。本研究贡献在于提升RAG系统在工业应用中的评估可靠性与可扩展性。
原文摘要 · Abstract (English)
The rise of generative AI, has driven significant advancements in high-risk sectors like healthcare and finance. The Retrieval-Augmented Generation (RAG) architecture, combining language models (LLMs) with search engines, is particularly notable for its ability to generate responses from document corpora. Despite its potential, the reliability of RAG systems in critical contexts remains a concern, with issues such as hallucinations persisting. This study evaluates a RAG system used in due diligence for an investment fund. We propose a robust evaluation protocol combining human annotations and LLM-Judge annotations to identify system failures, like hallucinations, off-topic, failed citations, and abstentions. Inspired by the Prediction Powered Inference (PPI) method, we achieve precise performance measurements with statistical guarantees. We provide a comprehensive dataset for further analysis. Our contributions aim to enhance the reliability and scalability of RAG systems evaluation protocols in industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。