让大模型学会在简单任务中假装不确定,提升安全场景下的可信度。
Inducing Artificial Uncertainty in Language Models
- 在无挑战数据时,人为制造虚假不确定性用于训练
- 用人工不确定数据训练的探测器,在难样本上校准效果更好
- 适合需要可靠置信度估计的安全敏感应用
在安全关键应用中,语言模型需能以有意义的概率表征其不确定性。许多不确定性量化方法依赖有监督数据,但对由海量网络爬取数据训练的大模型而言,找到合适的未见挑战性数据日益困难。若模型对所有预测都始终自信,则不确定性量化方法可能在新数据上持续高估置信度。因此,为高性能模型寻找足够不确定性的训练数据将愈发困难,且随大模型饱和数据集而加剧。为此,我们首次提出在语言模型中诱导人工不确定性的方法,并研究了在缺乏真实挑战数据时,如何在简单易题上引入人工不确定性。通过训练探测器识别原模型中的人工不确定性,发现基于人工不确定数据训练的探测器,在识别真实不确定性方面表现更优,尤其在难样本上校准性显著提升,同时在简单数据上性能损失极小。
原文摘要 · Abstract (English)
In safety-critical applications, language models should be able to characterize their uncertainty with meaningful probabilities. Many uncertainty quantification approaches require supervised data; however, finding suitable unseen challenging data is increasingly difficult for large language models trained on vast amounts of scraped data. If the model is consistently (and correctly) confident in its predictions, the uncertainty quantification method may consistently overestimate confidence on new and unfamiliar data. Finding data which exhibits enough uncertainty to train supervised uncertainty quantification methods for high-performance models may therefore be challenging, and will increase in difficulty as LLMs saturate datasets. To address this issue, we first introduce the problem of inducing artificial uncertainty in language models, then investigate methods of inducing artificial uncertainty on trivially easy data in the absence of challenging data at training time. We use probes trained to recognize artificial uncertainty on the original model, and find that these probes trained on artificial uncertainty outperform probes trained without artificial uncertainty in recognizing real uncertainty, achieving notably higher calibration on hard data with minimal loss of performance on easy data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。