测试大模型自我认知能力,发现其无法准确预判多领域表现,需外部干预才能安全使用。
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models
- 构建分层基准评估模型自我判断能力,覆盖8个实验、4个认知层级。
- 模型在多领域任务上自评误差高达0.500~0.943,普遍失效。
- 仅外部控制能降低错误率76%,说明靠架构约束比提升自知更有效。
我们提出MIRROR,一个包含八个实验、跨越四个元认知层级的基准,用于评估大语言模型是否能利用自我知识做出更好决策。在约25万次评估实例中,对来自8个实验室的16个模型进行了测试,使用五种独立的行为测量通道。核心实验覆盖全部模型;有特殊硬件需求的实验则明确标注适用模型子集。研究发现两个直接影响智能体部署的现象:(1) 组合式自我预测普遍失败——在原始15模型的Exp3-v1集上,组合校准误差为0.500至0.943(在扩展的16模型平衡版Exp3-v2上为0.434至0.758),表明模型无法预测其在跨领域任务中的表现;(2) 模型虽具备高于随机水平但不完美的领域特定自我认知,却系统性地无法将其转化为恰当的行动选择——外部元认知控制使自信失败率从0.600降至0.143(温度为0时降低76%,5个模型平均降低70%)。提供模型自身校准分数并未带来显著改善(p > 0.05);唯有架构约束有效。这表明,通往更安全自主AI系统的路径在于外部元认知支撑,而非提升自我认知。代码、数据及Croissant元数据将随基准公开。
原文摘要 · Abstract (English)
We introduce MIRROR, a benchmark comprising eight experiments across four metacognitive levels that evaluates whether large language models can use self-knowledge to make better decisions. We evaluate 16 models from 8 labs across approximately 250,000 evaluation instances using five independent behavioral measurement channels. Core experiments are run across the full model roster; experiments with specialized infrastructure requirements report explicitly marked model subsets. We find two phenomena with direct implications for agentic deployment: (1) compositional self-prediction fails universally -- the Compositional Calibration Error ranges from 0.500 to 0.943 on the original 15-model Exp3-v1 set (and 0.434 to 0.758 on the balanced 16-model Exp3-v2 expansion), indicating that models cannot predict their own performance on multi-domain tasks, and (2) models exhibit above-chance but imperfect domain-specific self-knowledge yet systematically fail to translate even this partial awareness into appropriate agentic action-selection -- external metacognitive control reduces the Confident Failure Rate from 0.600 to 0.143 (76% reduction at temperature 0; mean 70% at temperature 0.7 across 5 models from 4 labs). Providing models with their own calibration scores produces no significant improvement (p > 0.05); only architectural constraint is effective. This suggests that external metacognitive scaffolding -- not improved self-knowledge -- is the path to safer autonomous AI systems. Code, data, and Croissant metadata will be released publicly with the benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。