arXiv:2605.06673cs.CLcs.AI2026-05

33个前沿大模型在6个领域上的元认知监控能力差异显著,揭示了聚合评分的误导性。

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

论文配图:Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas
图 1 · 摘自论文原文
  • 按6个领域分组测试33个模型,用置信度评估元认知监控能力
  • 专业应用领域最易监控(平均AUROC 0.742),形式推理和自然科学最难
  • 模型家族内监控模式具一致性,适合用于部署前领域筛选

聚合的元认知质量评分掩盖了模型在MMLU基准不同领域间的内部差异。我们对来自8个模型家族的33个前沿大模型,针对6个预设领域的各250道题目(共1500题)进行了测试,并基于口头置信度(0-100)计算每个模型-领域单元的二型AUROC。总观测数达47,151次。所有在整体上表现优于随机水平的模型均展现出显著的领域级差异。应用/专业类知识是可监测性最高的领域(平均AUROC = 0.742),在33个模型中有21个位列前二;形式推理与自然科学则最难以监控,其中至少一个在27个模型中排名垫底。中间三个领域的表现统计上无显著差异(Kendall's W = 0.164)。领域内一致性分析(相似比=0.95)表明六领域分组是实用的基准分类体系,而非验证过的潜在结构。在Anthropic、Google-Gemini和Qwen家族中,模型族内监控轮廓形状聚类显著(置换检验p < .0001),而DeepSeek、Google-Gemma和OpenAI则不显著。Gemma 4 31B相较Gemma 3 27B提升+0.202 AUROC。三种在二元保留/撤回探测中被标记为无效的模型,在口头置信度下仍呈现正常监控特征,证明探测格式特异性。198个单元的95%置信区间中位宽度为0.199。总体分割稳定性(r = 0.893)高,但模型轮廓层面的分割稳定性较弱(中位r = 0.184)。结果表明,领域级差异在聚合指标下被掩盖,支持在特定应用场景部署前进行基准阶段的领域筛查。

原文摘要 · Abstract (English)

Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, under an a priori six-domain grouping) to 33 frontier LLMs from eight model families and computed Type-2 AUROC per model-domain cell using verbalized confidence (0-100). Total observations: 47,151. Every model with above-chance aggregate monitoring showed non-trivial domain-level variation. Applied/Professional knowledge was reliably the easiest benchmark domain to monitor (mean AUROC = .742, ranked top-2 in 21 of 33 models); Formal Reasoning and Natural Science were reliably the hardest (one of the two ranked bottom-2 in 27 of 33 models). The three middle domains were statistically indistinguishable (Kendall's W = .164). A subject-level coherence analysis (within-domain similarity ratio = 0.95) confirms the six-domain grouping is a pragmatic benchmark taxonomy, not a validated latent construct. Within-family profile-shape clustering is significant for Anthropic, Google-Gemini, and Qwen (permutation p < .0001) but not DeepSeek, Google-Gemma, or OpenAI. Gemma 4 31B showed a +.202 AUROC improvement over Gemma 3 27B. Three models classified Invalid on binary KEEP/WITHDRAW probes produced normal profiles under verbalized confidence, confirming probe-format specificity. Bootstrap 95% CIs on 198 cells have median width .199. Split-half aggregate stability r = .893; profile-level split-half is weaker (grand median r = .184). These results show stable benchmark-domain variation obscured by aggregate metrics, and support benchmark-stage domain screening as a step before deployment in specific application areas.

元认知大模型评估领域差异基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。