微调会降低语言模型置信度评分的可靠性,需谨慎使用。
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
- 分析置信度评分在微调后的变化机制
- 微调后置信度与输出质量相关性显著下降
- 提醒使用者必须验证置信度有效性
不确定性量化是衡量语言模型置信度的一类技术,可用于检测幻觉或提示用户审查不确定预测。要具备实用性,置信度评分必须与输出质量高度相关。然而近期研究表明,微调可能影响置信度评分与输出质量的相关性。本文研究置信度评分对监督微调(SFT)的敏感性,发现微调后多种置信度评分的相关性均下降,这可能源于置信度变化受输出与训练分布相似性等非质量因素影响。通过案例研究展示,若未解决这一误相关问题,将削弱置信度评分在下游任务中的实用性。研究结果表明,置信度指标不能直接套用,需验证其有效性,并推动开发更鲁棒的微调抗性度量方法。
原文摘要 · Abstract (English)
Uncertainty quantification is a set of techniques that measure confidence in language models. They can be used, for example, to detect hallucinations or alert users to review uncertain predictions. To be useful, these confidence scores must be correlated with the quality of the output. However, recent work found that fine-tuning can affect the correlation between confidence scores and quality. Hence, we investigate the underlying behavior of confidence scores to understand its sensitivity to supervised fine-tuning (SFT). We find that post-SFT, the correlation of various confidence scores degrades, which can stem from changes in confidence scores due to factors other than the output quality, such as the output's similarity to the training distribution. We demonstrate via a case study how failing to address this miscorrelation reduces the usefulness of the confidence scores on a downstream task. Our findings show how confidence metrics cannot be used off-the-shelf without testing, and motivate the need for developing metrics which are more robust to fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。