梳理大模型不确定性量化方法,提升高风险场景下的可信度。
Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey
- 提出按计算效率与不确定性维度分类的新框架
- 识别输入、推理、参数、预测四类独特不确定性源
- 适合关注大模型可靠性与安全性的研究者与从业者
大型语言模型(LLMs)在文本生成、推理和决策方面表现优异,已应用于医疗、法律和交通等高风险领域。然而,其输出常看似合理却错误,可靠性存疑。不确定性量化(UQ)通过估计输出置信度来增强可信性,支持风险控制和选择性预测。传统UQ方法因计算开销和解码不一致难以适配LLMs,且大模型引入了输入模糊性、推理路径分歧、解码随机性等新不确定性来源,超出经典似然性和认知不确定性范畴。为此,本文提出基于计算效率与不确定性维度(输入、推理、参数、预测)的新型分类体系,评估现有技术,分析实际应用潜力,并指出可扩展、可解释、鲁棒的UQ方法是提升大模型可靠性的关键挑战。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel in text generation, reasoning, and decision-making, enabling their adoption in high-stakes domains such as healthcare, law, and transportation. However, their reliability is a major concern, as they often produce plausible but incorrect responses. Uncertainty quantification (UQ) enhances trustworthiness by estimating confidence in outputs, enabling risk mitigation and selective prediction. However, traditional UQ methods struggle with LLMs due to computational constraints and decoding inconsistencies. Moreover, LLMs introduce unique uncertainty sources, such as input ambiguity, reasoning path divergence, and decoding stochasticity, that extend beyond classical aleatoric and epistemic uncertainty. To address this, we introduce a new taxonomy that categorizes UQ methods based on computational efficiency and uncertainty dimensions (input, reasoning, parameter, and prediction uncertainty). We evaluate existing techniques, assess their real-world applicability, and identify open challenges, emphasizing the need for scalable, interpretable, and robust UQ approaches to enhance LLM reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。