arXiv:2606.21937cs.CYcs.AI2026-06

让大模型更真实地评估自己,提升自评准确性。

Latent Confidence Alignment for LLM Self-Assessment

论文配图:Latent Confidence Alignment for LLM Self-Assessment
图 1 · 摘自论文原文
  • 基于潜在能力模型构建自评一致性度量方法
  • 在医学数据集上验证,自评质量显著提升
  • 适合关注模型可信度与推理效率的研究者

大语言模型的置信度校准通常通过比较预测置信度与实际准确率来评估。然而,此类方法未建模题目难度,难以解释偏差来源,也无法判断模型置信度是真实自省还是生成过程的副产品。为此,本文采用基于Rasch模型的潜在能力框架与元认知视角,提出潜伏置信度对齐误差(LCAE),用于衡量模型自评与由模型能力及题目难度决定的潜在错误概率之间的一致性。进一步引入题目难度作为外部信号,并结合推理机制进行优化。在包含20个模型的医学领域数据集上的实验表明,该方法在不损害模型能力的前提下,提升了自评质量,并揭示了可靠性与推理开销之间的关联。

原文摘要 · Abstract (English)

Confidence calibration in large language models (LLMs) is commonly evaluated by comparing predicted confidence with observed accuracy. However, such approaches do not model item difficulty, making it difficult to interpret discrepancies and to determine whether model confidence reflects genuine self-assessment or is merely a byproduct of the response generation process. To address this, we adopt a Rasch model-based latent ability framework and a metacognitive perspective, and propose Latent Confidence Alignment Error (LCAE) to measure the consistency between model self-assessment and the latent error probability implied by model ability and item difficulty. We further incorporate item difficulty as an external signal with a reasoning mechanism. Experiments on a medical-domain dataset with 20 models show that the proposed approach improves self-assessment quality without affecting model ability, and reveals an association between reliability and inference cost.

大模型自评置信度校准元认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。