用心理物理方法量化AI自我评估能力,帮系统更懂何时该自信、何时该求助。
Measuring the metacognition of AI
- 引入元判别力(meta-d')框架评估AI判断可靠性与置信度的区分能力。
- 基于信号检测理论,发现大模型在高风险场景下能自发降低决策置信度。
- 适用于需可信决策的AI系统,如医疗、自动驾驶等高风险领域。
稳健的决策过程必须考虑不确定性,尤其是在涉及固有风险时。随着人工智能系统越来越多地融入决策流程,管理不确定性越来越依赖于这些系统的元认知能力——即评估和调节自身决策可靠性的能力。因此,采用可靠的方法来衡量AI的元认知能力至关重要。本文主要提出方法论建议,主张采用元判别力(meta-d')框架作为评估AI元认知敏感性的黄金标准,即生成能区分正确与错误回答的置信度评分的能力。此外,我们提出利用信号检测理论(SDT)来衡量AI在不确定性和风险下自发调节决策的能力。为验证这些心理物理学框架的实际效用,我们在三个大型语言模型(LLMs)——GPT-5、DeepSeek-V3.2-Exp 和 Mistral-Medium-2508 上进行了两组实验。
原文摘要 · Abstract (English)
A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks. Because artificial intelligence (AI) systems are increasingly integrated into decision-making workflows, managing uncertainty relies more and more on the metacognitive capabilities of these systems; i.e, their ability to assess the reliability of and regulate their own decisions. Hence, it is crucial to employ robust methods to measure the metacognitive abilities of AI. This paper is primarily a methodological contribution arguing for the adoption of the meta-d' framework as the gold standard for assessing the metacognitive sensitivity of AIs--the ability to generate confidence ratings that distinguish correct from incorrect responses. Moreover, we propose to leverage signal detection theory (SDT) to measure the ability of AIs to spontaneously regulate their decisions based on uncertainty and risk. To demonstrate the practical utility of these psychophysical frameworks, we conduct two series of experiments on three large language models (LLMs)--GPT-5, DeepSeek-V3.2-Exp, and Mistral-Medium-2508.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。