arXiv:2504.18114cs.CLcs.AI2025-04EMNLP被引 15

评测6种幻觉检测指标,发现多数与人类判断不符。

Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection

  • 在4个数据集、37个模型上测试6类指标
  • 多数指标与人工判断不一致,参数量增大时效果不稳定
  • GPT-4评估和模式搜索解码法表现更优

幻觉严重影响语言模型的可靠性与应用推广,但其准确度量仍是一大挑战。尽管已有多种任务和领域特定的评估指标被提出用于衡量忠实性与事实性,但这些指标的鲁棒性和泛化能力尚未得到充分验证。本文对6类不同的幻觉检测指标进行了大规模实证评估,覆盖4个数据集、37个来自5个模型家族的语言模型以及5种解码方法。研究发现当前评估存在显著缺陷:指标常与人类判断不一致,对问题理解过于片面,且在参数规模增加时表现不一致。令人鼓舞的是,基于大语言模型的评估(尤其是GPT-4)整体表现最佳,模式搜索类解码方法在知识依赖场景下能有效减少幻觉。这些结果凸显了建立更稳健评估指标的必要性,以及改进幻觉缓解策略的重要性。

原文摘要 · Abstract (English)

Hallucinations pose a significant obstacle to the reliability and widespread adoption of language models, yet their accurate measurement remains a persistent challenge. While many task- and domain-specific metrics have been proposed to assess faithfulness and factuality concerns, the robustness and generalization of these metrics are still untested. In this paper, we conduct a large-scale empirical evaluation of 6 diverse sets of hallucination detection metrics across 4 datasets, 37 language models from 5 families, and 5 decoding methods. Our extensive investigation reveals concerning gaps in current hallucination evaluation: metrics often fail to align with human judgments, take an overtly myopic view of the problem, and show inconsistent gains with parameter scaling. Encouragingly, LLM-based evaluation, particularly with GPT-4, yields the best overall results, and mode-seeking decoding methods seem to reduce hallucinations, especially in knowledge-grounded settings. These findings underscore the need for more robust metrics to understand and quantify hallucinations, and better strategies to mitigate them.

幻觉检测评估指标大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。