用缺失信息量衡量大模型回答不确定性,发现熵比置信度更准。
LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information
- 将大模型视为隐式填补器,按缺失程度分级测试
- 熵随信息缺失上升,解释准确率方差能力比置信度强0.057
- 提出新诊断指标ρ_R(α),仅需重复采样即可评估上下文作用
大型语言模型在上下文不完整或退化场景中应用日益广泛。我们认为,在上下文不完整时生成答案的LLM可视为隐式填补器,并依据多重填补(MI)文献中的标准:不确定性应随缺失信息量增加而增大。我们在SQuAD上通过控制框架,将上下文可用性分为五个等级进行评估。考察了两种可从重复采样中估计的答案级不确定性度量:基于采样的置信度(经验模态频率)和响应熵。置信度未能反映缺失程度的增加:即使准确率崩溃,其值仍保持高位。而熵则随上下文移除逐渐升高,与MI类比一致,且在所有证据水平下对准确率方差的解释能力显著优于置信度(二次项R²差距高达0.057)。我们进一步提出黑箱诊断指标ρ_R(α),仅需带与不带上下文的重复采样,即可估计上下文级别α所解决的基线不确定性比例。结果表明,在上下文不完整条件下,熵是比置信度更敏感的黑箱不确定性度量。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in settings where the available context is incomplete or degraded. We argue that an LLM generating answers under incomplete context can be viewed as an implicit imputer, and evaluated against a criterion from the multiple imputation (MI) literature: uncertainty should scale with the amount of missing information. We assess this criterion on SQuAD, using a controlled framework in which context availability is varied across five levels. We evaluate two answer-level uncertainty measures that can be estimated from repeated sampling: sampling-based confidence (empirical mode frequency) and response entropy. Confidence fails to reflect increasing missingness: it remains high even as accuracy collapses. Entropy, by contrast, increases with context removal, consistent with the MI analogy, and explains substantially more variance in accuracy than confidence across all evidence levels (quadratic $R^2$ gap up to 0.057). We further introduce a black-box diagnostic $ρ_R(α)$ that estimates the proportion of baseline uncertainty resolved by context level $α$, requiring only repeated sampling with and without context. These results suggest that entropy is a more responsive black-box uncertainty measure than confidence under incomplete context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。