arXiv:2601.00348cs.CLcs.AI2026-01中稿 · IJCNN 2025被引 1

提出新方法提升大模型生成事实时的不确定性判断能力。

Robust Uncertainty Quantification for Factual Generation of Large Language Models

  • 设计包含虚假名称的陷阱问题,模拟复杂提问场景。
  • 在4个模型上测试,鲁棒性指标平均提升0.1-0.2(ROCAUC)。
  • 适合关注大模型可信生成、幻觉检测的研究者与开发者。

大语言模型(LLM)技术快速发展,已广泛应用于专业与日常生活,但其持续存在的幻觉问题严重削弱了生成内容的可靠性与可信度。当前不确定性量化方法虽在常规问答中表现良好,但在非标准或对抗性提问下效果显著下降,影响实际应用中的可信度。本文针对多事实生成任务,构建包含虚假名称的陷阱问题数据集,提出一种新型鲁棒不确定性量化方法(RU)。实验表明,该数据集具有优异的区分能力;在4个不同模型上对比基线方法,所提方法平均提升0.1-0.2的ROCAUC值,为解决大模型幻觉问题提供了新思路与有效工具。

原文摘要 · Abstract (English)

The rapid advancement of large language model(LLM) technology has facilitated its integration into various domains of professional and daily life. However, the persistent challenge of LLM hallucination has emerged as a critical limitation, significantly compromising the reliability and trustworthiness of AI-generated content. This challenge has garnered significant attention within the scientific community, prompting extensive research efforts in hallucination detection and mitigation strategies. Current methodological frameworks reveal a critical limitation: traditional uncertainty quantification approaches demonstrate effectiveness primarily within conventional question-answering paradigms, yet exhibit notable deficiencies when confronted with non-canonical or adversarial questioning strategies. This performance gap raises substantial concerns regarding the dependability of LLM responses in real-world applications requiring robust critical thinking capabilities. This study aims to fill this gap by proposing an uncertainty quantification scenario in the task of generating with multiple facts. We have meticulously constructed a set of trap questions contained with fake names. Based on this scenario, we innovatively propose a novel and robust uncertainty quantification method(RU). A series of experiments have been conducted to verify its effectiveness. The results show that the constructed set of trap questions performs excellently. Moreover, when compared with the baseline methods on four different models, our proposed method has demonstrated great performance, with an average increase of 0.1-0.2 in ROCAUC values compared to the best performing baseline method, providing new sights and methods for addressing the hallucination issue of LLMs.

大模型幻觉检测不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。