arXiv:2605.02915cs.CLcs.LG2026-05

自验证能提升模型置信度判断,但效果因任务和模型而异。

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

论文配图:When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal
图 1 · 摘自论文原文
  • 让模型自我审核答案,作为置信度信号。
  • 在ARC-Challenge上,部分模型自验证优于传统基线。
  • 适合特定任务与模型,不具普适性,需对比基线评估。

同模型自验证通过让模型自我审查预测结果,是一种有潜力的置信度信号,用于选择性预测。然而,当强基线(如LL-AVG和LL-SUM)被严肃考虑时,其实际价值尚不明确。我们在ARC-Challenge和TruthfulQA-MC数据集上,针对多种模型家族、规模及提示变体进行了评估,不仅考察正确性排序,还通过AURC和工作点分析衡量弃权质量。结果高度依赖任务与模型:在ARC-Challenge上,Phi-2和Qwen系列模型中,自验证显著优于LL-AVG,尤其在Qwen-7B上提升最大;但在TruthfulQA-MC上,信号可靠性较低,小模型易受提示影响,DeepSeek-R1-Distill-8B甚至劣于LL-AVG,且LL-SUM常为更强基线。因此,自验证不应视为通用不确定性估计器,而应理解为依赖任务类型、模型族、提示形式以及所比基线的条件性置信信号。

原文摘要 · Abstract (English)

Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its practical value remains unclear once strong likelihood-based baselines are taken seriously. We evaluate self-verification against two such baselines, LL-AVG and LL-SUM, on ARC-Challenge and TruthfulQA-MC across multiple model families, scales, and prompt variants. We measure not only correctness ranking, but also abstention quality through AURC and operating-point analyses. The results are sharply task- and model-dependent. On ARC-Challenge, self-verification substantially improves over LL-AVG for Phi-2 and the Qwen models, with the largest gains appearing in Qwen-7B. On TruthfulQA-MC, however, the signal is less reliable: smaller models can become prompt-sensitive, DeepSeek-R1-Distill-8B degrades relative to LL-AVG, and LL-SUM often remains the stronger practical baseline. We therefore do not treat self-verification as a general-purpose uncertainty estimator. In this setting, it is better understood as a conditional confidence signal whose value depends on task type, model family, prompt formulation, and, crucially, the baseline it must beat.

置信度估计自验证模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。