用模糊概率框架让大模型更真实地表达不确定性的高低层次。
Verbalizing LLM's Higher-order Uncertainty via Imprecise Probabilities
- 用提示词和后处理直接提取一阶与二阶不确定性
- 在模糊问答、上下文学习等场景中显著提升可信度
- 适合需要可靠置信度评估的决策系统
尽管对大语言模型(LLMs)不确定性提取的需求日益增长,但实证表明,基于经典概率框架的提取方法无法充分捕捉LLM的行为。这种不匹配导致在模糊问答、上下文学习和自我反思等场景中出现系统性失效。为此,我们提出基于模糊概率的新颖提示法,该框架能有效表征和提取高层次不确定性:一阶不确定性描述对响应结果的不确定,二阶不确定性则量化模型本身概率分布的模糊性。我们设计通用提示与后处理流程,可直接获取并量化这两层不确定性,并在多种任务中验证其有效性。该方法使LLM的不确定性报告更加忠实,提升可信度,支持下游决策。
原文摘要 · Abstract (English)
Despite the growing demand for eliciting uncertainty from large language models (LLMs), empirical evidence suggests that LLM behavior is not always adequately captured by the elicitation techniques developed under the classical probabilistic uncertainty framework. This mismatch leads to systematic failure modes, particularly in settings that involve ambiguous question-answering, in-context learning, and self-reflection. To address this, we propose novel prompt-based uncertainty elicitation techniques grounded in \emph{imprecise probabilities}, a principled framework for repesenting and eliciting higher-order uncertainty. Here, first-order uncertainty captures uncertainty over possible responses to a prompt, while second-order uncertainty (uncertainty about uncertainty) quantifies indeterminacy in the underlying probability model itself. We introduce general-purpose prompting and post-processing procedures to directly elicit and quantify both orders of uncertainty, and demonstrate their effectiveness across diverse settings. Our approach enables more faithful uncertainty reporting from LLMs, improving credibility and supporting downstream decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。