arXiv:2510.24222cs.CL2025-10被引 8

按知识与确定性双轴分类大模型幻觉,发现需针对性缓解策略。

HACK: Hallucinations Along Certainty and Knowledge Axes

  • 提出知识与确定性双轴分类框架,区分幻觉成因。
  • 验证不同幻觉类型在模型内知识与确定性上的差异。
  • 发现高确定性幻觉难被现有方法缓解,适合关注可信度的场景

大模型幻觉严重影响其可靠性。现有研究多基于外部特征分类,忽视内部机制差异,导致缓解策略不精准。本文提出沿知识与确定性双轴分类幻觉的框架,通过模型特定数据构建区分不同幻觉类型。在知识轴上,区分因缺知识导致的幻觉和虽有知识却仍幻觉的情况;通过控制激活的引导缓解方法验证两类差异显著。进一步分析显示,即使共享参数知识,模型间仍存在不同幻觉模式。在确定性轴上,识别出模型虽有正确知识却以高确定性幻觉的严重情形;引入新评估指标,发现部分缓解方法在平均表现良好,但对这类关键案例效果极差。结果表明,必须结合知识与确定性分析幻觉,并发展针对性缓解方案。

原文摘要 · Abstract (English)

Hallucinations in LLMs present a critical barrier to their reliable usage. Existing research usually categorizes hallucination by their external properties rather than by the LLMs' underlying internal properties. This external focus overlooks that hallucinations may require tailored mitigation strategies based on their underlying mechanism. We propose a framework for categorizing hallucinations along two axes: knowledge and certainty. Since parametric knowledge and certainty may vary across models, our categorization method involves a model-specific dataset construction process that differentiates between those types of hallucinations. Along the knowledge axis, we distinguish between hallucinations caused by a lack of knowledge and those occurring despite the model having the knowledge of the correct response. To validate our framework along the knowledge axis, we apply steering mitigation, which relies on the existence of parametric knowledge to manipulate model activations. This addresses the lack of existing methods to validate knowledge categorization by showing a significant difference between the two hallucination types. We further analyze the distinct knowledge and hallucination patterns between models, showing that different hallucinations do occur despite shared parametric knowledge. Turning to the certainty axis, we identify a particularly concerning subset of hallucinations where models hallucinate with certainty despite having the correct knowledge internally. We introduce a new evaluation metric to measure the effectiveness of mitigation methods on this subset, revealing that while some methods perform well on average, they fail disproportionately on these critical cases. Our findings highlight the importance of considering both knowledge and certainty in hallucination analysis and call for targeted mitigation approaches that consider the hallucination underlying factors.

幻觉分析大模型可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。