arXiv:2609.04539cs.CL2026-09被引 5

提升大模型置信度估计准确性,让系统知道何时该信任模型输出。

A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs

论文配图:A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs
图 1 · 摘自论文原文
  • 通过结构化推理与距离感知校准结合,全面评估模型置信度。
  • 在多个数据集上显著提升置信度估计的可靠性,尤其在对话和事实分类任务中表现优异。
  • 适合需要高可信度判断的场景,如医疗、金融等关键领域应用。

大型语言模型(LLMs)部署中的一个关键挑战是建立可靠的置信度估计机制,使系统能判断何时应信任模型输出,何时需寻求人工干预。本文提出一种校准反思方法(Calibrated Reflection),融合结构化推理与距离感知校准技术。该框架包含三项创新:(1) 最大置信度选择(MCS)方法,全面评估所有可能标签下的置信度;(2) 基于反思的提示机制,增强推理可靠性;(3) 距离感知校准技术,考虑标签间的序数关系。我们在HelpSteer2、Llama T-REx及一个专有对话数据集上进行评估,证明该方法在对话与事实分类任务中均具有效性。本工作推动了大模型置信度估计向更可靠、更校准的方向发展,助力模型信任与人工判断的明智决策。

原文摘要 · Abstract (English)

A critical challenge in deploying Large Language Models (LLMs) is developing reliable mechanisms to estimate their confidence, enabling systems to determine when to trust model outputs versus seek human intervention. We present a Calibrated Reflection approach for enhancing confidence estimation in LLMs, a framework that combines structured reasoning with distance-aware calibration technique. Our approach introduces three key innovations: (1) a Maximum Confidence Selection (MCS) method that comprehensively evaluates confidence across all possible labels, (2) a reflection-based prompting mechanism that enhances reasoning reliability, and (3) a distance-aware calibration technique that accounts for ordinal relationships between labels. We evaluate our framework on diverse datasets, including HelpSteer2, Llama T-REx, and a proprietary conversational dataset, demonstrating its effectiveness across both conversational and fact-based classification tasks. This work contributes to the broader goal of developing reliable and well-calibrated confidence estimation methods for LLMs, enabling informed decisions about model trust and human judgement.

置信度估计大模型校准推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。