现有幻觉检测评估方法严重失真,需改用更可靠的评判标准。
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
- 用人类判断标准重评检测模型,发现传统方法性能虚高。
- 部分方法在新评估下准确率下降高达45.9%。
- 简单长度规则竟可媲美复杂算法,暴露出评估体系缺陷。
大语言模型(LLMs)虽推动自然语言处理发展,但其幻觉问题严重影响部署可靠性。尽管已有众多幻觉检测方法,其评估多依赖基于词项重叠的ROUGE指标,该指标与人类判断严重不符。通过全面的人类实验,我们发现虽然ROUGE具有高召回率,但精度极低,导致性能估计严重失真。实际上,多个成熟检测方法在采用更贴近人类判断的LLM-as-Judge等语义对齐指标后,性能下降最高达45.9%。分析还显示,基于响应长度的简单启发式规则可媲美复杂检测技术,暴露当前评估范式的根本缺陷。我们主张采用语义感知且稳健的评估框架,才能真实衡量幻觉检测效果,确保大模型输出可信。
原文摘要 · Abstract (English)
Large language models (LLMs) have revolutionized natural language processing, yet their tendency to hallucinate poses serious challenges for reliable deployment. Despite numerous hallucination detection methods, their evaluations often rely on ROUGE, a metric based on lexical overlap that misaligns with human judgments. Through comprehensive human studies, we demonstrate that while ROUGE exhibits high recall, its extremely low precision leads to misleading performance estimates. In fact, several established detection methods show performance drops of up to 45.9\% when assessed using human-aligned metrics like LLM-as-Judge. Moreover, our analysis reveals that simple heuristics based on response length can rival complex detection techniques, exposing a fundamental flaw in current evaluation practices. We argue that adopting semantically aware and robust evaluation frameworks is essential to accurately gauge the true performance of hallucination detection methods, ultimately ensuring the trustworthiness of LLM outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。