为大模型设计跨领域自监控评估工具,揭示其自信与纠错能力的三种模式。
The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring

- 基于人类元认知理论,设计524个跨领域任务测试模型自我监控。
- 20个前沿模型中,准确率与自监控敏感度呈反向关系,三类行为模式清晰可分。
- 适合研究大模型可信度、安全评估及智能系统可解释性的研究人员。
我们提出一个跨领域的大型语言模型自监控行为测试集,基于Nelson和Narens(1990)的元认知框架,并借鉴人类心理测量方法评估大模型。该测试集包含6个认知领域(学习、元认知校准、社会认知、注意力、执行功能、前瞻调控)共524个任务,每项均源自经典实验范式。任务T1-T5在OSF上预先注册,T6为探索性扩展。每次选择后,采用Koriat和Goldsmith(1996)的双探针法,要求模型决定是否保留或撤回答案、是否下注或放弃。关键指标为“撤回差值”:错误与正确题目的撤回率之差。对20个前沿大模型(共10,480次评估)的应用显示,模型呈现三种符合元认知架构的模式:普遍自信、普遍撤回、选择性敏感。准确率排名与元认知敏感度排名基本呈倒置关系。回顾性监控与前瞻性调控可分离(r = .17,95%CI较宽,样本量n=20;主要证据来自示例分析)。元认知校准的规模效应因模型架构而异:单调递减(Qwen)、单调递增(GPT-5.4)或持平(Gemma)。行为结果与独立的类型2信号检测理论方法结构一致,初步验证了跨方法构念效度。所有任务、数据与代码已公开:https://github.com/synthiumjp/metacognitive-monitoring-battery。
原文摘要 · Abstract (English)
We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework and applying human psychometric methodology to LLM evaluation. The battery comprises 524 items across six cognitive domains (learning, metacognitive calibration, social cognition, attention, executive function, prospective regulation), each grounded in an established experimental paradigm. Tasks T1-T5 were pre-registered on OSF prior to data collection; T6 was added as an exploratory extension. After every forced-choice response, dual probes adapted from Koriat and Goldsmith (1996) ask the model to KEEP or WITHDRAW its answer and to BET or decline. The critical metric is the withdraw delta: the difference in withdrawal rate between incorrect and correct items. Applied to 20 frontier LLMs (10,480 evaluations), the battery discriminates three profiles consistent with the Nelson-Narens architecture: blanket confidence, blanket withdrawal, and selective sensitivity. Accuracy rank and metacognitive sensitivity rank are largely inverted. Retrospective monitoring and prospective regulation appear dissociable (r = .17, 95% CI wide given n=20; exemplar-based evidence is the primary support). Scaling on metacognitive calibration is architecture-dependent: monotonically decreasing (Qwen), monotonically increasing (GPT-5.4), or flat (Gemma). Behavioural findings converge structurally with an independent Type-2 SDT approach, providing preliminary cross-method construct validity. All items, data, and code: https://github.com/synthiumjp/metacognitive-monitoring-battery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。