用嵌入距离分析检测大模型幻觉,方法更可靠。
Hallucination Detection: A Probabilistic Framework Using Embeddings Distance Analysis
- 通过嵌入空间的闵可夫斯基距离分析,发现幻觉内容有结构性差异。
- 在多种参数下,幻觉响应的嵌入距离分布显著不同,准确率达66%。
- 适合研究大模型可信性、幻觉检测或系统验证的开发者参考。
幻觉是影响大语言模型广泛应用的主要问题。现有检测方法多依赖启发式规则,本文提出一种数学严谨的推理框架,首次证明幻觉内容在嵌入空间中具有与正确内容不同的结构特征。基于闵可夫斯基距离分析,实验显示幻觉与真实内容的嵌入距离分布存在统计显著差异,且该差异在不同距离范数及关键词、问题或回复数量下均保持不变(尺度无关)。利用这一结构差异,构建了幻觉检测工具,在特定参数配置下达到66%的准确率,与领域内最佳结果相当。该方法具有新颖性和潜力,为后续研究提供了新方向。
原文摘要 · Abstract (English)
Hallucinations are one of the major issues affecting LLMs, hindering their wide adoption in production systems. While current research solutions for detecting hallucinations are mainly based on heuristics, in this paper we introduce a mathematically sound methodology to reason about hallucination, and leverage it to build a tool to detect hallucinations. To the best of our knowledge, we are the first to show that hallucinated content has structural differences with respect to correct content. To prove this result, we resort to the Minkowski distances in the embedding space. Our findings demonstrate statistically significant differences in the embedding distance distributions, that are also scale free -- they qualitatively hold regardless of the distance norm used and the number of keywords, questions, or responses. We leverage these structural differences to develop a tool to detect hallucinated responses, achieving an accuracy of 66\% for a specific configuration of system parameters -- comparable with the best results in the field. In conclusion, the suggested methodology is promising and novel, possibly paving the way for further research in the domain, also along the directions highlighted in our future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。