arXiv:2511.00620cs.CL2025-11

大模型输出的确定性与理论概率不一致,影响可信决策

Certain but not Probable? Differentiating Certainty from Probability in LLM Token Outputs for Probabilistic Scenarios

  • 用令牌概率和熵值评估模型确定性,对比理论分布
  • 模型准确率虽达100%,但概率预测偏离理论值
  • 适合关注大模型可信度与决策支持的开发者

可靠的不确定性量化对确保大语言模型在决策支持等知识密集型应用中的可信使用至关重要。模型确定性可从令牌逻辑值估算,衍生的概率与熵值能反映其在任务上的表现。然而,在概率场景中,该方法可能不足,因令牌输出概率应与理论结果概率对齐。本文研究了在明确概率场景下,令牌确定性与理论概率分布的一致性。使用GPT-4.1与DeepSeek-Chat,评估十种涉及概率的提示(如掷六面骰),包含有无显式概率提示(如掷公平六面骰)。测量两个维度:(1)响应是否符合场景约束;(2)令牌级输出概率与理论概率的对齐程度。结果表明,尽管两种模型在所有场景中均达到100%域内准确率,但其令牌级概率与熵值始终偏离理论分布。

原文摘要 · Abstract (English)

Reliable uncertainty quantification (UQ) is essential for ensuring trustworthy downstream use of large language models, especially when they are deployed in decision-support and other knowledge-intensive applications. Model certainty can be estimated from token logits, with derived probability and entropy values offering insight into performance on the prompt task. However, this approach may be inadequate for probabilistic scenarios, where the probabilities of token outputs are expected to align with the theoretical probabilities of the possible outcomes. We investigate the relationship between token certainty and alignment with theoretical probability distributions in well-defined probabilistic scenarios. Using GPT-4.1 and DeepSeek-Chat, we evaluate model responses to ten prompts involving probability (e.g., roll a six-sided die), both with and without explicit probability cues in the prompt (e.g., roll a fair six-sided die). We measure two dimensions: (1) response validity with respect to scenario constraints, and (2) alignment between token-level output probabilities and theoretical probabilities. Our results indicate that, while both models achieve perfect in-domain response accuracy across all prompt scenarios, their token-level probability and entropy values consistently diverge from the corresponding theoretical distributions.

不确定性量化大模型可信度概率推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。