arXiv:2512.10451cs.LG2025-12中稿 · NeurIPS被引 2

让模型学会评估自己,动态选择最可信的预测者

Metacognitive Sensitivity for Test-Time Dynamic Model Selection

  • 用心理测量学指标meta-d'衡量模型自我判断的可靠性
  • 在多个数据集上提升集成模型的整体准确率
  • 适合需要高可靠性的AI决策场景

人类认知的关键特征是元认知——评估自身知识与判断可靠性。尽管深度学习模型能表达置信度,但常存在校准不良的问题,即所表达的置信度无法反映真实能力。模型真的知道它知道什么吗?借鉴人类认知科学,我们提出一种评估和利用人工智能元认知的新框架。引入基于心理学的元认知敏感性度量meta-d',用于刻画模型置信度预测自身准确性的可靠性。随后,将这一动态敏感性分数作为上下文,由基于强化学习的仲裁器执行测试时模型选择,学习在给定任务中信任哪个专家模型。在多个数据集及深度学习模型组合(包括CNN和视觉语言模型)上的实验表明,该元认知方法可提升集成推理的准确率,超越各组成模型。本工作为人工智能模型提供了一种新的行为解释,将集成选择重构为同时评估短期信号(置信度预测分数)与中期特质(元认知敏感性)的问题。

原文摘要 · Abstract (English)

A key aspect of human cognition is metacognition - the ability to assess one's own knowledge and judgment reliability. While deep learning models can express confidence in their predictions, they often suffer from poor calibration, a cognitive bias where expressed confidence does not reflect true competence. Do models truly know what they know? Drawing from human cognitive science, we propose a new framework for evaluating and leveraging AI metacognition. We introduce meta-d', a psychologically-grounded measure of metacognitive sensitivity, to characterise how reliably a model's confidence predicts its own accuracy. We then use this dynamic sensitivity score as context for a bandit-based arbiter that performs test-time model selection, learning which of several expert models to trust for a given task. Our experiments across multiple datasets and deep learning model combinations (including CNNs and VLMs) demonstrate that this metacognitive approach improves joint-inference accuracy over constituent models. This work provides a novel behavioural account of AI models, recasting ensemble selection as a problem of evaluating both short-term signals (confidence prediction scores) and medium-term traits (metacognitive sensitivity).

元认知模型选择置信度校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。