arXiv:2604.17716cs.CLcs.AI2026-04被引 2

用三类可信度信号筛选模型,有效提升精准预测能力。

Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction

论文配图:Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction
图 1 · 摘自论文原文
  • 将大模型置信度分为有效、不确定、无效三类进行评估
  • 有效模型在选择性预测中平均准确率高出近27个百分点
  • 适合关注模型可靠性与决策安全性的研究者

有效性筛查(Cacioli, 2026d, 2026e)将大模型的置信度信号分类为有效、不确定或无效。我们检验此类别是否能预测选择性预测的表现。对来自七个系列的20个前沿大模型,在六个认知任务共524个题目上进行了评估。有效模型的平均二类AUROC为.624(标准差.048),无效模型为.357(标准差.231),效应量Cohen's d = 2.81,p = .002。三类层级单调递增:无效(.357)< 不确定(.554)< 有效(.624)。分半交叉验证显示中位效应量d = 1.77,P(d > 0) = 1.0,覆盖1000次分割。该三分类体系解释了AUROC方差的47%。DeepSeek-R1在全覆盖率下准确率为85.3%,降至10%覆盖率时骤降至11.3%。筛查结果可预测实际表现,对选择性预测具有关键意义。

原文摘要 · Abstract (English)

The validity screen (Cacioli, 2026d, 2026e) classifies LLM confidence signals as Valid, Indeterminate, or Invalid. We test whether these classifications predict selective prediction performance. Twenty frontier LLMs from seven families were evaluated on 524 items across six cognitive tracks. Valid models show mean Type 2 AUROC = .624 (SD = .048). Invalid models show mean AUROC = .357 (SD = .231). Cohen's d = 2.81, p = .002. The tiers order monotonically: Invalid (.357) < Indeterminate (.554) < Valid (.624). Split-half cross-validation yields median d = 1.77, P(d > 0) = 1.0 across 1,000 splits. The three-tier classification accounts for 47% of the variance in AUROC. DeepSeek-R1 drops from 85.3% accuracy at full coverage to 11.3% at 10% coverage. The screen predicts the criterion. For selective prediction, the screen matters.

大模型可靠性置信度评估选择性预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。