arXiv:2503.21157cs.LG2025-03被引 3

对比五种无需参考答案的幻觉检测模型,评估其在RAG中的表现。

Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best?

  • 采用无参考基准方法,直接判断大模型生成内容是否幻觉。
  • 在六类RAG应用中,部分模型达到高精度与高召回率。
  • 适合关注生成可信度、需自动化评估的研究者和开发者。

本文调研了用于自动检测检索增强生成(RAG)中幻觉的评估模型,并在六种RAG应用场景中全面评测了它们的性能。研究涵盖的方法包括:LLM-as-a-Judge、Prometheus、Lynx、Hughes幻觉评估模型(HHEM)以及可信语言模型(TLM)。这些方法均为无参考型,无需真实答案或标注即可识别错误的LLM响应。研究表明,在多种RAG应用中,部分方法能以高精度和高召回率持续检测出不正确的RAG输出。

原文摘要 · Abstract (English)

This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. Methods included in our study include: LLM-as-a-Judge, Prometheus, Lynx, the Hughes Hallucination Evaluation Model (HHEM), and the Trustworthy Language Model (TLM). These approaches are all reference-free, requiring no ground-truth answers/labels to catch incorrect LLM responses. Our study reveals that, across diverse RAG applications, some of these approaches consistently detect incorrect RAG responses with high precision/recall.

RAG幻觉检测评估模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。