发现大模型内部隐藏更多纠错信息,可精准定位错误类型。
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- 识别出特定令牌内蕴含真相信息,提升检测准确率
- 模型内部能预判错误类型,支持定制化修正策略
- 内部正确答案与外部输出不一致,揭示认知偏差
大语言模型常产生事实错误、偏见和推理失败,统称为“幻觉”。近期研究发现,模型内部状态编码了输出真实性的信息,可用于错误检测。本文首次发现,真实性信息高度集中在特定令牌中,利用该特性显著提升错误检测性能。然而,此类检测器在跨数据集上表现不佳,表明真实性编码并非普遍通用,而是多维度的。进一步发现,内部表示可用于预测模型可能产生的错误类型,有助于制定针对性缓解策略。最后揭示一个关键现象:模型内部虽编码了正确答案,却仍持续生成错误输出。这些发现从模型内部视角深化了对幻觉的理解,为未来错误分析与缓解研究提供新方向。
原文摘要 · Abstract (English)
Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have demonstrated that LLMs' internal states encode information regarding the truthfulness of their outputs, and that this information can be utilized to detect errors. In this work, we show that the internal representations of LLMs encode much more information about truthfulness than previously recognized. We first discover that the truthfulness information is concentrated in specific tokens, and leveraging this property significantly enhances error detection performance. Yet, we show that such error detectors fail to generalize across datasets, implying that -- contrary to prior claims -- truthfulness encoding is not universal but rather multifaceted. Next, we show that internal representations can also be used for predicting the types of errors the model is likely to make, facilitating the development of tailored mitigation strategies. Lastly, we reveal a discrepancy between LLMs' internal encoding and external behavior: they may encode the correct answer, yet consistently generate an incorrect one. Taken together, these insights deepen our understanding of LLM errors from the model's internal perspective, which can guide future research on enhancing error analysis and mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。