arXiv:2509.24988cs.CLcs.AI2025-09被引 5

不靠模型自省,用历史预测模式提升大模型自信度估计的准确性。

Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns

  • 基于历史输出模式构建通用正确性模型,无需依赖模型自我判断。
  • 在5个模型家族、2个数据集上验证,信心估计准确率显著提升。
  • 适合需要可靠置信度评估的高风险应用场景,如医疗与金融决策。

为高风险或面向用户的应用部署大语言模型,生成准确且校准良好的置信度估计至关重要,但仍是开放挑战。以往研究常将置信度视为模型‘自我认知’问题,即大模型能否判断自身回答是否正确;该假设隐含模型自身可获取答案正确性的专属信息。然而实验发现,大模型预测自身输出正确性表现与无关模型无异。我们提出一种通用正确性模型(GCM),关键在于引入目标模型的历史预测数据。通过多种方法注入历史正确性信息,发现GCM可在多个大模型和数据集上学习通用正确性模式。进一步分析表明,答案表述方式是正确性的重要预测因子。探索非训练式历史注入方法发现,将历史作为上下文示例可提升预测效果,事后校准也能有效降低校准误差。在Qwen3-8B上评估,覆盖5个模型家族、MMLU与TriviaQA数据集及下游选择性预测任务,结果表明可靠的置信度估计是一种可泛化、模型无关的技能,其核心是系统编码正确性历史,而非依赖自我内省。

原文摘要 · Abstract (English)

Generating accurate and calibrated confidence estimates is critical for deploying LLMs in high-stakes or user-facing applications, and remains an open challenge. Prior research has often framed confidence as a problem of eliciting a model's "self-knowledge", i.e., the ability of an LLM to judge whether its own answers are correct; this approach implicitly assumes that there is some privileged information about the answer's correctness that is accessible to the model itself. However, our experiments reveal that an LLM attempting to predict the correctness of its own outputs generally performs no better than an unrelated LLM. Moreover, we hypothesize that a key factor in building a "Correctness Model" (CM) is exposure to a target model's historical predictions. We propose multiple methods to inject this historical correctness information, creating a Generalized Correctness Model (GCM). We first show that GCMs can be trained on the correctness data from many LLMs and learn patterns for correctness prediction applicable across datasets and models. We then use CMs as a lens for studying the source of correctness prediction ability and its generalization, systematically controlling their training data and finding that answer phrasing is a strong predictor for correctness. We further explore alternative methods of injecting history without training an LLM, finding that including history as in-context examples can help improve correctness prediction, and post-hoc calibration can provide complementary reductions in calibration error. We evaluate GCMs based on Qwen3-8B across 5 model families and the MMLU and TriviaQA datasets, as well as on a downstream selective prediction task, finding that reliable LLM confidence estimation is a generalizable and model-agnostic skill learned by systematically encoding correctness history rather than a model-specific skill reliant on self-introspection.

置信度估计大模型通用模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。