检测大模型自信与真实能力的差距,发现同一模型在不同任务上表现不一。
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
- 设计五项任务,拆解大模型自信行为为五个维度
- 发现Gemini 2.5 Flash在任务间自信度差异达47分
- 适合关注模型可靠性与可信推理的研究者
Metacognitive Probe 是一个包含五项任务、十五个槽位的探索性诊断工具,将大语言模型的自信行为分解为五个行为上可区分的维度:信心校准(T1-CC)、认知警觉性(T2-EV)、知识边界(T3-KB)、校准范围(T4-CR)和推理链验证(T5-RCV)。该工具在 N=8 个前沿模型和 N=69 名人类中进行评估。其灵感来自 Flavell(1979)和 Nelson 与 Narens(1990),但基于可观测的自信-正确性对齐,而非跨物种元认知量表,且预设的人类发展假设被证伪。综合基准测试(MMLU、BIG-Bench、HELM、GPQA)仅判断模型回答是否正确,却无法揭示模型是否知道自己答错。模型可能在整体校准得分80分,但在局部仍极度高估。Metacognitive Probe 能揭示这些隐藏弱点。主要发现:Gemini 2.5 Flash 在同一模型内存在47分的跨任务自信差异——任务内校准最佳(T1-CC = 88;Spearman rho = +0.551,95% CI [+0.14, +0.80],p = 0.005),而跨任务难度预测最差(T4-CR = 41;sigma_conf = 1.4,覆盖十二个事实点)。
原文摘要 · Abstract (English)
The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct dimensions: confidence calibration (T1-CC), epistemic vigilance (T2-EV), knowledge boundary (T3-KB), calibration range (T4-CR), and reasoning-chain validation (T5-RCV). It is evaluated on N=8 frontier models and N=69 humans. The instrument is motivated by Flavell (1979) and Nelson and Narens (1990) but operates on observable confidence-correctness alignment; it is not a validated cross-species metacognition scale, and the pre-specified human developmental hypothesis was falsified. Composite benchmarks (MMLU, BIG-Bench, HELM, GPQA) ask whether a model produces a correct response. They are silent on whether the model knows when its response is wrong. A model can score 80 on a composite calibration benchmark and still be wildly overconfident in narrow pockets the aggregate cannot surface. The Metacognitive Probe surfaces those pockets. Our headline is a 47-point within-model dissociation in Gemini 2.5 Flash: panel-best within-task calibration (T1-CC = 88; Spearman rho = +0.551, 95% CI [+0.14, +0.80], p = 0.005) and panel-worst cross-task difficulty prediction (T4-CR = 41; sigma_conf = 1.4 across twelve factoids).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。