发现大模型的言语不确定性可线性调控,有效降低幻觉
Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations
- 用线性特征建模言语不确定性,与语义不确定性的相关性较弱
- 言语与语义不确定性不一致是更优的幻觉预测指标
- 推理时干预该特征,短回答幻觉平均减少约30%
大语言模型在做出错误陈述时常表现出过度自信的言辞风格,这种‘过度自信的幻觉’误导用户并削弱信任。我们发现,‘言语不确定性’在模型表示空间中由单一线性特征决定,且与模型实际的‘语义不确定性’相关性较弱。基于此,我们证明:(1) 言语与语义不确定性之间的差异比语义不确定性本身更能预测幻觉;(2) 可在推理阶段对言语不确定性进行干预,从而在短文本回答中减少自信型幻觉,实现平均约30%的相对降幅。
原文摘要 · Abstract (English)
LLMs often adopt an assertive language style also when making false claims. Such ``overconfident hallucinations'' mislead users and erode trust. Achieving the ability to express in language the actual degree of uncertainty around a claim is therefore of great importance. We find that ``verbal uncertainty'' is governed by a single linear feature in the representation space of LLMs, and show that this has only moderate correlation with the actual ``semantic uncertainty'' of the model. We apply this insight and show that (1) the mismatch between semantic and verbal uncertainty is a better predictor of hallucinations than semantic uncertainty alone and (2) we can intervene on verbal uncertainty at inference time and reduce confident hallucinations on short-form answers, achieving an average relative reduction of ~30%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。