为大模型生成的不确定性提供系统性分类与评估框架。
The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

- 按输入、参数、分词、解码四个层面拆解大模型不确定性来源。
- 共识类方法在多个任务中表现最优,大模型规模越大不确定性越低。
- 适用于大模型可信生成、安全部署及方法对比的实用诊断工具。
大语言模型(LLM)虽在推理和内容生成方面取得进展,但其固有的随机性严重影响预测可靠性。传统不确定性分类(如认知与偶然不确定性)难以刻画生成过程的多阶段特性,也难评估各类不确定性量化(UQ)方法的有效性。本文提出细粒度不确定性分类体系,将LLM不确定性分解为输入级、参数级、标记级和解码过程级四类,并将现有UQ方法归为贝叶斯、集成、共识型与单次遍历四类。我们构建涵盖多种生成场景与指标的综合评估框架,在TriviaQA、GSM8K和HumanEval等基准上,对Qwen3、Llama 3.2和DeepSeek-V3三类主流模型中的21种典型UQ方法进行实证评估。结果表明:(i) UQ方法效果受任务类型与生成设置影响显著;(ii) 共识型方法(Deg与EigV)持续优于其他方法;(iii) 模型规模越大,不确定性估计值越低,揭示了大模型不确定性存在经验尺度规律。本研究弥合理论溯源与实际应用之间的鸿沟,提供一套系统化的大模型不确定性诊断工具。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods. In this paper, we propose a granular uncertainty taxonomy that systematically attributes LLM uncertainty into input-level, parameter-level, token-level, and decoding-process sources. Correspondingly, we categorize existing UQ methods into Bayesian, ensemble, consensus-based, and single-pass approaches. Furthermore, we introduce a comprehensive evaluation framework covering diverse generation settings and metrics. We empirically evaluate 21 typical UQ methods across three prominent LLM families, including Qwen3, Llama 3.2, and DeepSeek-V3, on benchmarks such as TriviaQA, GSM8K, and HumanEval. Our experimental results demonstrate that (i) the effectiveness of UQ methods is sensitive to task types and generation settings; (ii) consensus-based methods, typed Deg and EigV, consistently outperform other UQ approaches; and (iii) larger model scales correlate with lower uncertainty estimates, suggesting an empirical scaling law for LLM uncertainty. This work bridges the gap between theoretical origins and practical deployment, providing a versatile diagnostic tool for systematically quantifying uncertainty in LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。