arXiv:2605.04295cs.LGcs.AI2026-05中稿 · publication in the…

通过自适应语义熵量化大模型输出不确定性,提升安全关键场景可靠性。

LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy

论文配图:LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy
图 1 · 摘自论文原文
  • 基于多响应聚类的语义熵,动态评估提示级不确定性。
  • 在TriviaQA上实现0.88的AUROC,优于传统词元熵方法(0.65)。
  • 结合置信校准,提供有限样本下误差率可控的决策保障。

大模型的过度自信,尤其是在幻觉时,给安全关键场景的应用带来挑战,因此可靠地估计不确定性至关重要。现有方法多依赖词汇或概率度量,但常忽略语义相近却表述不同的响应之间的差异。本文提出自适应共形语义熵(ACSE),通过自适应测量大模型对同一提示生成的多种响应的语义分散程度,实现提示级不确定性估计。该评分函数基于多个多样化响应的聚类语义熵,并根据每个簇的语义特征动态调整得分。为确保统计可靠性,采用共形校准设定接受/拒绝决策规则,在有限样本且不依赖分布假设的前提下,保证被接受响应的错误率不超过用户指定容忍值。大量实验在不同大模型与数据集上验证,本方法在判别性能、共形保证和概率校准指标上均显著优于当前最优基线。特别地,在TriviaQA数据集上,本方法的AUROC达0.88,远超词元熵方法的0.65。

原文摘要 · Abstract (English)

LLMs' overconfidence, particularly when hallucinating, poses a significant challenge for the deployment of the models in safety-critical settings and makes a reliable estimation of uncertainty necessary. Existing approaches for uncertainty quantification typically prioritize lexical or probabilistic measures; however, these techniques often ignore the semantic variance of different responses with similar meaning. In this paper, we propose Adaptive Conformal Semantic Entropy (ACSE), a method for estimating prompt-level uncertainty by adaptively measuring semantic dispersion in LLMs outputs. Our uncertainty scoring function is based on clustering semantic entropy of multiple diverse responses to the same prompt. The function adaptively adjusts the uncertainty score based on semantic features of each cluster. To ensure statistical reliability of our score, we use conformal calibration to apply a decision rule to accept/abstain the prompts, providing a finite-sample, distribution-free guarantee such that the error rate among the accepted responses remains bounded by a user-specified tolerance. Our extensive experimental evaluations using different LLMs and datasets, demonstrate that our approach consistently outperforms state-of-the-art uncertainty quantification baselines using discriminative performance, conformal guarantees, and probabilistic calibration indicators. As a highlight, for TriviaQA dataset, AUROC of our approach is 0.88 compared to 0.65 produced by the token entropy approach.

不确定性量化大模型共形预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。