arXiv:2604.24070cs.CLcs.AI2026-04

让小模型学会说真话:用自洽性信号训练,提升回答可信度。

Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B

  • 用自洽性生成可信度标签,训练小模型输出更真实的自信判断。
  • 在TriviaQA上,可信度判别准确率达77.4%,接近自洽性信号的99.9%。
  • 模型需正确标签保持输出格式,否则会陷入盲目自信陷阱。

小型指令微调大模型在极简提示下会产生退化的语言可信度:置信度天花板超95%,类型2 AUROC接近随机水平,有效性分布无效。本文测试了以自洽性推导的目标进行信心条件监督微调(CSFT)能否弥合内部信息与口头读出之间的差距。预注册的第0阶段实验在Gemma 3 4B上进行,仅使用正确主答案项训练,结果为负面:类型2 AUROC从0.554降至0.509,源于训练目标中标签熵坍塌。探索性救援移除了筛选器,对全部2,000个校准样本进行训练。该方法在保留10样本自洽信号(类型2 AUROC = 0.999)的同时,实现单次读出超越对数熵(0.701),在独立验证集TriviaQA上达到类型2 AUROC = 0.774。随机目标对照组无提升(0.501)。在MMLU上,准确率从54.2%提升至77.4%(基线为56.1%),支持目标依赖性解释。结果具有探索性,为二值而非连续校准,且仅在单一规模下观察到。研究提出两点设计教训:信心训练需保持标签熵,正确目标能正则化输出格式。

原文摘要 · Abstract (English)

Small instruct-tuned LLMs produce degenerate verbal confidence under minimal elicitation: ceiling rates above 95%, near-chance Type-2 AUROC, and Invalid validity profiles. We test whether confidence-conditioned supervised fine-tuning (CSFT) with self-consistency-derived targets can close the gap between internal information and verbal readout. A pre-registered Phase 0 protocol on Gemma 3 4B-it with a modal filter restricting training to items with correct modal answers produced a negative result: AUROC2 dropped from 0.554 to 0.509 due to label-entropy collapse in the training targets. An exploratory rescue removed the filter, training on all 2,000 calibration items. This produced a binary verbal correctness discriminator with AUROC2 = 0.774 on held-out TriviaQA, compressing a 10-sample self-consistency signal (AUROC2 = 0.999) into a single-pass readout exceeding logit entropy (0.701). The shuffled-target control showed no improvement (0.501). On MMLU, accuracy improved from 54.2% to 77.4% with the shuffled model at baseline (56.1%), supporting a target-dependent interpretation. The result is exploratory, binary rather than continuously calibrated, and observed at a single scale. It identifies two design lessons: confidence training requires label entropy, and correct targets regularise output format.

可信度小模型自洽性微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。