arXiv:2608.13430cs.CLcs.AI2026-08

指令微调让模型更自信,但理性表述的词汇多样性却变差了。

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

  • 通过对比基础模型与指令微调模型,发现信心提升但准确率变化小。
  • 指令微调使推理文本的跨样本多样性下降,表面词汇多样性则不一致。
  • 即使控制答案选择和长度,信心与多样性仍独立变化,说明二者本质不同。

指令微调的语言模型在多种生成任务中表现优异,但近期研究发现其存在言语性过度自信现象。在问答任务中,模型的过度自信可能与生成推理过程的一致性相关。本文研究指令微调是否引起生成答案推理文本的词汇多样性变化。我们在多个问答基准上评估三对匹配的基线模型与指令微调模型,发现指令微调显著提升模型自信度,尽管预测准确率提升有限,且基于似然的概率校准能力下降。其次,指令微调对推理多样性的影响不均匀:跨推理文本的多样性持续下降,而表面词汇多样性在不同模型和基准间呈现方向与幅度各异的变化。最后,即使控制答案选择和推理长度后,这种差异依然存在,证实模型信心与推理多样性是指令微调带来的两种独立效应。

原文摘要 · Abstract (English)

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.

语言模型指令微调信心评估词汇多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。