新基准T RIVIA+填补了长上下文与噪声标签的空白,助力更真实地评估大模型幻觉检测。
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights

- 构建基于RAG的长上下文幻觉检测基准T RIVIA+,支持真实场景测试。
- 首次提供四类不同噪声模式的标签,模拟现实中的标注误差。
- 发现现有检测器仍有提升空间,且噪声显著影响检测效果。
幻觉指大模型生成的不忠实、虚构或不一致内容,影响广泛。现有幻觉检测基准(HDBs)普遍存在缺陷:缺乏基于RAG的长上下文数据集(因长度限制人工标注困难),且未提供真实世界中常见的标签噪声。本文提出一个理想基准应具备的属性,并构建开源的T RIVIA+基准,其样本上下文长度为文献中最长,且包含四组具有样本依赖和独立噪声模式的标注数据。在多个SOTA检测器上测试发现:当前检测器在基于RAG的基准上仍有巨大提升空间;基础的LLM-as-a-Judge方法表现优异;标签噪声会显著降低检测性能。研究结果与基准将推动面向RAG任务的幻觉检测研究。
原文摘要 · Abstract (English)
Hallucination, broadly referring to unfaithful, fabricated, or inconsistent content generated by LLMs, has wide-ranging implications. Therefore, a large body of effort has been devoted to detecting LLM hallucinations, as well as designing benchmark datasets for evaluating these detectors. In this work, we first establish a desiderata of properties for hallucination detection benchmarks (HDBs) to exhibit for effective evaluation. A critical look at existing HDBs through the lens of our desiderata reveals that none of them exhibits all the properties. We identify two largest gaps: (1) RAG-based grounded benchmarks with long context are severely lacking (partly because length impedes human annotation); and (2) Existing benchmarks do not make available realistic label noise for stress-testing detectors although real-world use-cases often grapple with label noise due to human or automated/weak annotation. To close these gaps, we build and open-source a new RAG-based HDB called T RIVIA+ that underwent a rigorous human annotation process. Notably, our benchmark exhibits all desirable properties including (1) T RIVIA+ contains samples with the longest context in the literature; and (2) we design and share four sets of noisy labels with different, both sample-dependent and sampleindependent, noise schemes. Finally, we perform experiments on RAG-based HDBs, including our T RIVIA+, using popular SOTA detectors that reveal new insights: (i) ample room remains for current detectors to reach the performance ceiling on RAG-based HDBs, (ii) the basic LLM-as-a-Judge baseline performs competitively, and (iii) label noise hinders detection performance. We expect that our findings, along with our proposed benchmark 1 , will motivate and foster needed research on hallucination detection for RAG-based tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。