首个跨领域通用幻觉检测基准,覆盖真实场景多文档输出。
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
- 构建涵盖多领域、多格式的无任务依赖幻觉数据集
- 长文本上下文下检测准确率下降明显,影响RAG系统可靠性
- 开源模型中Quotient Detections表现最优,F1达0.84
随着大语言模型在高风险领域的广泛应用,检测非基于证据生成的幻觉内容已成为关键挑战。现有幻觉检测基准多为合成数据,聚焦于抽取式问答,难以反映真实场景中多文档上下文与完整句子输出的复杂性。本文提出HalluMix基准,一个多样化、任务无关的数据集,涵盖多个领域和格式。利用该基准,我们评估了七种幻觉检测系统(含开源与闭源),揭示了不同任务、文档长度和输入表示下的性能差异。分析显示,长上下文场景下性能显著下降,对实际检索增强生成(RAG)部署有重要影响。其中Quotient Detections表现最佳,准确率为0.82,F1得分为0.84。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed in high-stakes domains, detecting hallucinated content$\unicode{x2013}$text that is not grounded in supporting evidence$\unicode{x2013}$has become a critical challenge. Existing benchmarks for hallucination detection are often synthetically generated, narrowly focused on extractive question answering, and fail to capture the complexity of real-world scenarios involving multi-document contexts and full-sentence outputs. We introduce the HalluMix Benchmark, a diverse, task-agnostic dataset that includes examples from a range of domains and formats. Using this benchmark, we evaluate seven hallucination detection systems$\unicode{x2013}$both open and closed source$\unicode{x2013}$highlighting differences in performance across tasks, document lengths, and input representations. Our analysis highlights substantial performance disparities between short and long contexts, with critical implications for real-world Retrieval Augmented Generation (RAG) implementations. Quotient Detections achieves the best overall performance, with an accuracy of 0.82 and an F1 score of 0.84.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。