研究大模型如何用语言标记表达真实可信度,发现其表现仍不靠谱。
Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence?

- 提出内生置信度指标(MIC),量化模型对语言标记的信任程度。
- 不同模型和任务下,语言标记的置信度差异不显著,稳定性差。
- 适合关注大模型可信赖性与输出一致性研究者阅读。
大模型的语言自信表达应忠实反映其内在不确定性。尽管已有研究表明大模型在使用表征性标记(如“可能…”)时难以符合人类认知,但尚不清楚模型能否在其自身的语言自信体系中,将特定标记与稳定的置信水平关联起来,以及上下文特征如何影响这一能力。本文首次系统性地研究该问题,将「标记内生置信度」(MIC)定义为模型在特定任务领域中对某一表征性标记所关联的估计内在置信度。我们提出了7项指标,评估MIC在分布内与跨分布的稳定性。应用于多种模型与任务后发现,即使在以模型为中心解释标记意义的情况下,大模型依然存在显著的校准偏差,难以在不同分布间区分标记的置信水平,尽管其在任务间保持了相对一致的排序。这为理解大模型的忠实校准提供了关键、互补的证据,强调需改进标记使用的对齐性与稳定性,以提升可信度与可靠性。
原文摘要 · Abstract (English)
LLMs' linguistically expressed confidence should faithfully reflect their intrinsic uncertainty. While recent work shows LLMs struggle to use epistemic markers (e.g., "it is likely...") in a human-aligned fashion, it remains unclear whether models can apply their own linguistic confidence framework to associate markers with specific confidence levels in a stable and generalizable way, and how contextual features impact this ability. We conduct the first systematic study of this question, formalizing _marker internal confidence_ (MIC) as the estimated intrinsic confidence a model associates with a specific epistemic marker in a given task domain. We present 7 metrics to evaluate the stability of MICs within and across distributions. Applying our analysis framework to diverse models and tasks, we find that LLMs remain faithfully miscalibrated even under model-centric interpretation of marker meanings, struggling to differentiate markers by internal confidence across distributions despite preserving a somewhat consistent ranking order across tasks. This supplies critical, complementary evidence to existing work toward a holistic understanding of faithful calibration in LLMs, emphasizing the need for more aligned and stable marker use to improve trustworthiness and reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。