arXiv:2507.22915cs.CLcs.AI2025-07被引 9

为大模型幻觉提供理论框架与应对方案

Theoretical Foundations and Mitigation of Hallucination in Large Language Models

  • 区分内在与外在幻觉,定义幻觉风险并给出理论边界
  • 提出统一检测与缓解流程,涵盖检索增强、校准等方法
  • 给出评估幻觉的标准化数据集和指标,适合研究者参考

大语言模型中的幻觉指生成内容与输入或真实世界事实不符。本文对幻觉进行了严谨的理论分析,包括形式化定义与学习理论框架下的风险边界推导(PAC-Bayes与Rademacher复杂度)。区分了内在与外在幻觉,定义了模型的幻觉风险,并推导其理论上限。系统梳理了幻觉检测策略,如标记级不确定性估计、置信度校准与注意力对齐检查。在缓解方面,讨论了检索增强生成、幻觉感知微调、逻辑校准及事实验证模块等方法。提出一个统一的检测与缓解工作流,并附示意图。最后,建议使用特定数据集、指标与实验设计来量化与降低幻觉。本工作为解决大模型幻觉问题提供了理论基础与实践指南。

原文摘要 · Abstract (English)

Hallucination in Large Language Models (LLMs) refers to the generation of content that is not faithful to the input or the real-world facts. This paper provides a rigorous treatment of hallucination in LLMs, including formal definitions and theoretical analyses. We distinguish between intrinsic and extrinsic hallucinations, and define a \textit{hallucination risk} for models. We derive bounds on this risk using learning-theoretic frameworks (PAC-Bayes and Rademacher complexity). We then survey detection strategies for hallucinations, such as token-level uncertainty estimation, confidence calibration, and attention alignment checks. On the mitigation side, we discuss approaches including retrieval-augmented generation, hallucination-aware fine-tuning, logit calibration, and the incorporation of fact-verification modules. We propose a unified detection and mitigation workflow, illustrated with a diagram, to integrate these strategies. Finally, we outline evaluation protocols for hallucination, recommending datasets, metrics, and experimental setups to quantify and reduce hallucinations. Our work lays a theoretical foundation and practical guidelines for addressing the crucial challenge of hallucination in LLMs.

大模型幻觉理论分析检测缓解评估协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。