arXiv:2604.22215cs.CLcs.AI2026-04被引 2

7个3-9B大模型的口头自信表达无效,无法真实反映不确定性。

Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen

  • 用数值和分类两种方式测试模型自信表达,均采用贪婪解码。
  • 7个模型在数值自信上全部失效,平均天花板率达91.7%。
  • 推理长度越长,自信越低,适合做心理测量有效性筛查。

口头自信提取被广泛用于获取大语言模型的不确定性估计。本研究在预注册实验(OSF: osf.io/azbvx)中,对七个指令微调的开源模型(3-9B参数,四个家族)在最小数值提示下使用贪婪解码,测试其在项目级二级区分上的最低有效标准。524个TriviaQA题目在消费级硬件上以Q5_K_M量化运行,产生8,384次确定性试验。对每个模型-格式单元进行心理测量有效性筛查。所有七种指令模型在数值自信上均被判定为无效(H2确认,7/7对比预测≥4/7),平均天花板率高达91.7%(H1确认)。分类提示未能改善有效性,反而导致六个模型任务准确率低于5%(H4未确认)。在观测方差范围内,标记级对数概率无法有效预测口头自信(H5确认,平均交叉验证R² < 0.01)。在推理精炼模型中,推理轨迹长度与自信呈强负相关(rho = -0.36, p < .001),符合推理污染效应。结果不意味着内部不确定性表示不存在,而是表明在该参数规模下,最小口头提示无法保留在输出接口处的内部信号。心理测量筛查应在下游使用此类信号前执行。

原文摘要 · Abstract (English)

Verbal confidence elicitation is widely used to extract uncertainty estimates from LLMs. We tested whether seven instruction-tuned open-weight models (3-9B parameters, four families) produce verbalised confidence that meets minimal validity criteria for item-level Type-2 discrimination under minimal numeric elicitation with greedy decoding. In a pre-registered study (OSF: osf.io/azbvx), 524 TriviaQA items were administered under numeric (0-100) and categorical (10-class) elicitation to eight models at Q5_K_M quantisation on consumer hardware, yielding 8,384 deterministic trials. A psychometric validity screen was applied to each model-format cell. All seven instruct models were classified Invalid on numeric confidence (H2 confirmed, 7/7 vs. predicted >=4/7), with a mean ceiling rate of 91.7% (H1 confirmed). Categorical elicitation did not rescue validity. Instead, it disrupted task performance in six of seven models, producing accuracy below 5% (H4 not confirmed). Token-level logprobability did not usefully predict verbalised confidence under the observed variance regime (H5 confirmed, mean cross-validated R^2 < 0.01). Within the reasoning-distilled model, reasoning-trace length showed a strong negative partial correlation with confidence (rho = -0.36, p < .001), consistent with the Reasoning Contamination Effect. These results do not imply that internal uncertainty representations are absent. They show that minimal verbal elicitation fails to preserve internal signals at the output interface in this model-size regime. Psychometric screening should precede any downstream use of such signals.

大模型不确定性心理测量可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。