arXiv:2502.08109cs.CLcs.AI2025-02被引 11

提出HuDEx模型,让大模型不仅能识别幻觉,还能解释原因,提升可信度。

HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses

  • 将幻觉检测与解释能力结合,实现错误自检
  • 在幻觉检测上超越Llama3 70B和GPT-4
  • 零样本和多数据集测试均表现稳定,适合高精度场景

大语言模型在自然语言处理任务中虽取得显著进展,但幻觉问题仍严重影响其可靠性,尤其在要求高事实准确性的领域。现有评估基准多聚焦于幻觉检测与事实性判断,缺乏解释能力。本文提出一种增强解释性的幻觉检测模型HuDEx,通过整合检测与解释机制,使用户及模型自身能理解并减少错误。实验表明,该模型在幻觉检测准确率上优于Llama3 70B与GPT-4,且保持可靠解释能力。在零样本及其他测试环境中均表现良好,展现出跨数据集的强适应性。该方法为幻觉检测研究引入可解释性融合新范式,显著提升评估可靠性。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have shown promising improvements, often surpassing existing methods across a wide range of downstream tasks in natural language processing. However, these models still face challenges, which may hinder their practical applicability. For example, the phenomenon of hallucination is known to compromise the reliability of LLMs, especially in fields that demand high factual precision. Current benchmarks primarily focus on hallucination detection and factuality evaluation but do not extend beyond identification. This paper proposes an explanation enhanced hallucination-detection model, coined as HuDEx, aimed at enhancing the reliability of LLM-generated responses by both detecting hallucinations and providing detailed explanations. The proposed model provides a novel approach to integrate detection with explanations, and enable both users and the LLM itself to understand and reduce errors. Our measurement results demonstrate that the proposed model surpasses larger LLMs, such as Llama3 70B and GPT-4, in hallucination detection accuracy, while maintaining reliable explanations. Furthermore, the proposed model performs well in both zero-shot and other test environments, showcasing its adaptability across diverse benchmark datasets. The proposed approach further enhances the hallucination detection research by introducing a novel approach to integrating interpretability with hallucination detection, which further enhances the performance and reliability of evaluating hallucinations in language models.

幻觉检测可解释AI大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。