让大模型从单一信心判断,转向多维度自我评估。
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

- 提出六维评估框架,分解模型自我认知为努力、能力等维度。
- 努力维度在推理任务中表现最优,且不随模型大小过度乐观。
- 适合关注模型可靠性与安全部署的研究者和工程师。
大型语言模型(LLMs)在需要可靠自我评估的场景中日益重要。传统基于概率的正确性估计已逐渐被口头表达的信心取代,但信心被证明是不一致且过于乐观的预测指标。基于人类心理学中的认知评价理论,我们提出一种多维模型自我评估视角,引入六个基于评价的评估维度(除信心外),并在12个LLM和38项跨越八个领域的任务上评估其预测模型失败的能力。结果表明,与能力相关的评价维度(尤其是努力和能力)在多数场景下表现优于或匹配信心;其中,努力维度不仅更准确,且在不同模型规模下保持稳定,不过度乐观。情感维度仅提供微弱预测信号。此外,最有效的评估维度随任务特性系统性变化:推理密集型任务中努力维度预测力最强,而检索类任务中能力与信心主导。总体而言,结构化的多维自我评估是提升大模型在多样化真实场景中部署可靠性和安全性的有效路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved from using probabilistic correctness estimates to, more recently, eliciting verbalized confidence. Confidence, however, has been shown to be an inconsistent and overoptimistic predictor of model correctness. Drawing on cognitive appraisal theory, a framework from human psychology that decomposes self-evaluation into multiple components, we propose a multidimensional perspective on model self-assessment. We elicit six appraisal-based dimensions of self-assessment, alongside confidence, and evaluate their utility for predicting model failure across 12 LLMs and 38 tasks spanning eight domains. We find that competence-related appraisal dimensions, particularly effort and ability, consistently match or outperform confidence across most settings. Effort additionally yields less overoptimistic estimates that remain stable across model sizes. In contrast, affective dimensions provide marginally predictive signals. Furthermore, the most informative dimension varies systematically with task characteristics: effort is most predictive for reasoning-intensive tasks, while ability and confidence dominate on retrieval-oriented tasks. Broadly, our findings indicate that structured multidimensional self-assessment is a promising approach to improving the reliability and safety of language model deployment across diverse real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。