LLM不确定度量化本质是无监督聚类,无法识别自信幻觉。
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

- 将模型生成一致性误作不确定性,实为聚类算法
- 高置信度错误答案仍被判定为低风险,忽略事实正确性
- 适合关注模型可靠性与评估方法的AI研究者
不确定性量化(UQ)被视为大语言模型在高风险场景部署的关键保障。然而我们指出,该领域存在根本性分类错误:当前主流方法本质上是无监督聚类算法。这些方法衡量的是模型生成内部的一致性,而非外部事实正确性。因此,它们无法检测‘自信幻觉’——即模型对稳定但错误的答案表现出高置信度。这导致现有方法可能在部署时制造虚假的安全感。具体而言,这种依赖内部状态导致三个关键问题:超参数敏感性危机使部署不可靠,内部评估循环将稳定性误当作真实性,且缺乏真实标签迫使使用不稳定的代理指标评估不确定性。为此,我们倡导范式转变,提出研究社区应采用更优评估指标与设置,实现原生不确定性机制,并以客观真理为锚点,确保模型置信度真正反映现实。
原文摘要 · Abstract (English)
Uncertainty Quantification (UQ) is widely regarded as the primary safeguard for deploying Large Language Models (LLMs) in high-stakes domains. However, we argue that the field suffers from a category error: mainstream UQ methods for LLMs are just unsupervised clustering algorithms. We demonstrate that most current approaches inherently quantify the internal consistency of the model's generations rather than their external correctness. Consequently, current methods are fundamentally blind to factual reality and fail to detect ``confident hallucinations,'' where models exhibit high confidence in stable but incorrect answers. Therefore, the current UQ methods may create a deceptive sense of safety when deploying the models with uncertainty. In detail, we identify three critical pathologies resulting from this dependence on internal state: a hyperparameter sensitivity crisis that renders deployment unsafe, an internal evaluation cycle that conflates stability with truth, and a fundamental lack of ground truth that forces reliance on unstable proxy metrics to evaluate uncertainty. To resolve this impasse, we advocate for a paradigm shift to UQ and outline a roadmap for the research community to adopt better evaluation metrics and settings, implement mechanism changes for native uncertainty, and anchor verification in objective truth, ensuring that model confidence serves as a reliable proxy for reality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。