用信号检测理论拆分大模型的真知与自知,发现自知能力不随准确率变化。
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
- 将分词概率视为信心信号,用信号检测理论区分知识量与自知力
- 不同模型自知力差1.98倍,且在科技类问题最弱
- 自知力能精准预测放弃回答后的性能提升,适合评估可信推理
标准的大模型置信度评估常混淆两种能力:模型真正知道多少(类型1准确率)和其信心信号是否真实反映知识(类型2元认知敏感性)。本文引入信号检测理论,将分词级归一化对数概率作为连续信心变量,答案正确性作为待判别状态。通过分析元认知受试者工作特征曲线(z-ROC)的非等方差结构,提出无需依赖双选项任务的模型无关信息度量——归一化元认知信息(meta-I_2r)。在224,000次事实问答测试中发现:(1) 元认知信息在不同模型间差异达1.98倍,且与准确率无关(TriviaQA相关系数-0.80,Natural Questions为+0.00);(2) 信心信号具有模型特异的非等方差结构(z-ROC斜率0.78–1.18),校准指标无法捕捉;(3) 元认知效率呈领域特异性,科技类问题最差;(4) 温度升高使准确率下降但元认知信息基本不变(三模型保持平稳);(5) 元认知信息与基于信心弃答的准确率提升完全正相关(rho=+1.00),而准确率本身不具此预测力。所有估计均经置换检验与自助法置信区间验证。本版本修正了自动评分器的长度偏差,经1,830次人工核验确认;此前报告的准确率-效率负耦合关系在重新标注后不再成立。
原文摘要 · Abstract (English)
Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows (Type-1 accuracy) and how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity). We apply Signal Detection Theory to decompose them, treating token-level normalised log-probability as a graded confidence variable and answer correctness as the state to be discriminated. We characterise the Type-2 ROC of this signal, including its unequal-variance structure via z-ROC analysis, and -- because the meta-d' efficiency ratio is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision -- quantify efficiency with a model-free information measure, normalised metacognitive information (meta-I_2r). Across four LLMs and 224,000 factual QA trials we find: (1) metacognitive information varies by a factor of 1.98 across models and is not predicted by accuracy, the rank correlation being -0.80 on TriviaQA and +0.00 on Natural Questions; (2) the confidence signal has model-specific unequal-variance structure (z-ROC slopes 0.78 to 1.18) invisible to calibration metrics, the slope ordering replicating on NQ; (3) efficiency is domain-specific, weakest in Science & Technology for every model; (4) temperature dissociates accuracy from metacognitive information, which stays near-flat for three of four models while accuracy falls; and (5) metacognitive information tracks the accuracy gain from confidence-based abstention exactly (rho = +1.00) while accuracy does not. All estimates carry permutation nulls and bootstrap confidence intervals. This version (v3) corrects a differential length bias in the automated correctness scorer, validated against 1,830 human adjudications; the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive relabelling. See the version note on page 1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。