大模型自信程度不反映真实能力,更像是阈值差异而非自我认知。
LLMs Show No Signs Of Individuated Metacognition

- 通过因子分析发现模型自信主要受共同难度轴影响,非个体化元认知。
- 去掉所有模型一致的题目后,自信与表现的相关性完全消失。
- 数学推理中的异常是因模型用推理过程代替自我评估,非真正自知。
我们对20个前沿大语言模型在六个基准上的二元自信判断进行分解,采用四分相关因子分析结合成对校准。在事实回忆和信息检索任务中,跨模型自信矩阵近似秩一,单一主导因子解释了大部分潜在方差。模型共享项目级难度轴,仅在决策阈值上不同。当移除所有模型一致的题目后,自信与表现的关系彻底崩溃。即使统计显著的模型间校准也极小,控制基线差异后几乎消失。数学推理看似例外,实为模型通过思维链解题来回答自信问题,绕过了我们想测量的子符号自我认知。在所有测试领域均未发现显著的显式个体化元认知证据。
原文摘要 · Abstract (English)
Confidence-weighted routing, selective abstention, and ensemble weighting all assume that a model's stated confidence is informative about its capability on the question being asked. They presume functional metacognition, the capacity to assess one's own capabilities, without exercising them. Aggregate calibration is well studied, with mixed results, but the underlying structure of elicited confidence is less well understood. We decompose binary confidence judgements from 20 frontier Large Language Models (LLMs) across six benchmarks using tetrachoric factor analysis paired with pairwise calibration, asking whether two models that differ in confidence also differ in performance. On factual recall and information retrieval benchmarks the cross-model confidence matrix is approximately rank-one and a single dominant factor captures most of the latent variance. Models retrieving facts share an item-level difficulty axis and differ mainly in their decision thresholds along it. Across all benchmarks the relationship between confidence and performance collapses once items that all models agree on are removed. Inter-model pairwise calibration is small even where statistically significant, and what remains shrinks to nothing once base-rate differences along the shared factor are controlled for. Mathematical reasoning is the apparent exception, but this turns out to be a confound where reasoning models answer questions about their confidence by trying to solve them in their chain of thought, bypassing the sub-symbolic self-knowledge we seek to measure. We find no evidence for significant verbalised individuated metacognition in any tested domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。