arXiv:2502.11677cs.CL2025-02ACL被引 27

用大模型内部状态提升对知识边界的感知能力

Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception

  • 利用模型生成前的内部状态预判自身信心,节省算力
  • 新方法C³使未知识别率在NQ上提升5.6%、HotpotQA上4.9%
  • 适合关注模型可靠性与安全性的应用开发者

大语言模型在多任务上表现优异,但常无法准确判断自身知识边界,导致自信地给出错误答案。本文从效率与风险角度探索利用模型内部状态增强其知识边界感知能力。研究发现,大模型可在生成前基于内部状态估算置信度,且该预判能力在自然问题(Natural Questions)、HotpotQA和MMLU等数据集上表现出显著效果,生成前后感知差距保持稳定。为降低关键领域风险,提出基于置信度一致性的校准方法 $C^3$,通过重述问题评估置信度一致性。实验显示,$C^3$ 显著提升模型识别知识盲区的能力,在NQ上未知感知率提高5.6%,在HotpotQA上提高4.9%。结果表明,生成前置信度估计可优化效率,而 $C^3$ 能有效控制输出风险,提升模型在实际应用中的可靠性。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit impressive performance across diverse tasks but often struggle to accurately gauge their knowledge boundaries, leading to confident yet incorrect responses. This paper explores leveraging LLMs' internal states to enhance their perception of knowledge boundaries from efficiency and risk perspectives. We investigate whether LLMs can estimate their confidence using internal states before response generation, potentially saving computational resources. Our experiments on datasets like Natural Questions, HotpotQA, and MMLU reveal that LLMs demonstrate significant pre-generation perception, which is further refined post-generation, with perception gaps remaining stable across varying conditions. To mitigate risks in critical domains, we introduce Confidence Consistency-based Calibration ($C^3$), which assesses confidence consistency through question reformulation. $C^3$ significantly improves LLMs' ability to recognize their knowledge gaps, enhancing the unknown perception rate by 5.6% on NQ and 4.9% on HotpotQA. Our findings suggest that pre-generation confidence estimation can optimize efficiency, while $C^3$ effectively controls output risks, advancing the reliability of LLMs in practical applications.

大模型知识边界置信度校准可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。