arXiv:2409.10102cs.IRcs.AI2024-09综述被引 128

系统评估RAG模型在六大维度的可信度,发现性能差距并提出评测基准。

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey

论文配图:Trustworthiness in Retrieval-Augmented Generation Systems: A Survey
图 1 · 摘自论文原文
  • 构建信任度评估框架Trust-RAG Compass,覆盖事实性、鲁棒性等六维
  • 设计TRC Bench评测基准,测试多种开源与闭源模型表现差异
  • 揭示不同模型在可信度维度上的短板,为实际应用提供指导

检索增强生成(RAG)已成为大语言模型发展的重要范式。尽管现有研究多关注准确性和效率,但对RAG系统的可信度仍缺乏深入探讨。RAG可通过引入外部实时知识提升模型可靠性,减少幻觉。然而,检索不可靠或知识利用不当仍可能导致不良输出。为此,本文提出统一框架Trust-RAG Compass,从事实性、鲁棒性、公平性、透明性、可问责性和隐私性六个维度评估RAG系统的可信度。在此框架下,系统梳理各维度现有研究。进一步,构建评测基准TRC Bench,对多种开源与闭源模型进行综合评估。结果揭示了不同类型大模型在可信度各维度上的性能差距。最后,基于发现提出关键挑战与未来研究方向。本工作旨在为后续研究提供结构化基础,并为真实场景中可信RAG系统开发提供实践指引。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) has quickly grown into a pivotal paradigm in the development of Large Language Models (LLMs). Although existing research mainly emphasizes accuracy and efficiency, the trustworthiness of RAG systems remains insufficiently explored. RAG can improve LLM reliability by grounding responses in external and up-to-date knowledge, reducing hallucinations. However, unreliable retrieval or improper knowledge utilization may still lead to undesirable outputs. To address these concerns, we propose a unified framework, Trust-RAG Compass, that assesses the trustworthiness of RAG systems across six key dimensions: factuality, robustness, fairness, transparency, accountability, and privacy. Within this framework, we provide a thorough review of the existing literature along each dimension. Furthermore, we introduce an evaluation benchmark, TRC Bench (\underline{T}rust-\underline{R}AG \underline{C}ompass \underline{Bench}mark), regarding the six dimensions and conduct comprehensive evaluations for a variety of proprietary and open-source models. Our results shed light on the performance gaps between different types of LLMs across varying dimensions of trustworthiness. Finally, we identify key challenges and promising directions for future research based on our findings. Through this work, we aim to provide a structured foundation for subsequent investigations and practical guidance for developing trustworthy RAG systems in real-world scenarios.

可信度RAG评测基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。