arXiv:2505.04847cs.CLcs.AI2025-05EMNLP被引 29

用新评估框架提升大模型在检索增强生成中的可信度测试

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

  • 引入FaithJudge评估框架,用人类标注样例提升自动幻觉检测
  • 新排行榜显示大模型在摘要等任务中仍存在高幻觉率
  • 适合关注生成式AI可靠性与评测标准的研究者和开发者

检索增强生成(RAG)旨在通过外部上下文降低大语言模型(LLMs)的幻觉问题,但即使提供相关上下文,模型仍常引入无支持信息或矛盾内容。本文介绍两项互补工作:一是自2023年起使用HHEM幻觉检测模型追踪各模型幻觉率的原始幻觉排行榜;二是针对现有检测方法局限性,提出基于LLM作为裁判的FaithJudge框架,利用多样化人工标注的幻觉样本显著提升自动化评估能力。新增强型幻觉排行榜聚焦于摘要、问答和数据到文本生成任务中,系统评估大模型在RAG中的忠实度。FaithJudge为大模型幻觉评估提供了更可靠的基准,有助于构建更可信的生成式AI系统。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) aims to reduce hallucinations by grounding responses in external context, yet large language models (LLMs) still frequently introduce unsupported information or contradictions even when provided with relevant context. This paper presents two complementary efforts at Vectara to measure and benchmark LLM faithfulness in RAG. First, we describe our original hallucination leaderboard, which has tracked hallucination rates for LLMs since 2023 using our HHEM hallucination detection model. Motivated by limitations observed in current hallucination detection methods, we introduce FaithJudge, an LLM-as-a-judge framework that leverages a pool of diverse human-annotated hallucination examples to substantially improve the automated hallucination evaluation of LLMs. We introduce an enhanced hallucination leaderboard centered on FaithJudge that benchmarks LLMs on RAG faithfulness in summarization, question-answering, and data-to-text generation tasks. FaithJudge enables a more reliable benchmarking of LLM hallucinations in RAG and supports the development of more trustworthy generative AI systems: https://github.com/vectara/FaithJudge.

大模型评估幻觉检测RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。