arXiv:2602.09158cs.LGcs.AI2026-02被引 1

揭示几何幻觉检测指标实际捕捉的是哪种幻觉类型

What do Geometric Hallucination Detection Metrics Actually Measure?

  • 构建可控制幻觉特性的合成数据集,分离不同幻觉属性
  • 发现不同几何统计量针对不同幻觉类型(如无关、不连贯)
  • 提出归一化方法提升跨领域检测性能,多领域AUROC提升34点

幻觉仍是生成模型在高风险应用中部署的障碍,尤其在缺乏外部真实答案验证时。为此,研究者关注大语言模型内部状态中的几何信号,这些信号能预测幻觉且仅需少量外部知识。由于幻觉可能由多种因素引起(如无关或不连贯),本文探讨现有几何检测指标究竟捕捉了幻觉的哪些具体特性。为此,我们构建了一个合成数据集,系统地改变输出的正确性、置信度、相关性、连贯性和完整性等属性。结果表明,不同的几何统计量对不同类型的幻觉具有敏感性。同时发现,许多现有方法对任务领域变化(如数学题与历史题)极为敏感。为此,我们提出一种简单归一化方法,有效缓解领域偏移影响,在多领域设置下实现AUROC提升34个百分点。

原文摘要 · Abstract (English)

Hallucination remains a barrier to deploying generative models in high-consequence applications. This is especially true in cases where external ground truth is not readily available to validate model outputs. This situation has motivated the study of geometric signals in the internal state of an LLM that are predictive of hallucination and require limited external knowledge. Given that there are a range of factors that can lead model output to be called a hallucination (e.g., irrelevance vs incoherence), in this paper we ask what specific properties of a hallucination these geometric statistics actually capture. To assess this, we generate a synthetic dataset which varies distinct properties of output associated with hallucination. This includes output correctness, confidence, relevance, coherence, and completeness. We find that different geometric statistics capture different types of hallucinations. Along the way we show that many existing geometric detection methods have substantial sensitivity to shifts in task domain (e.g., math questions vs. history questions). Motivated by this, we introduce a simple normalization method to mitigate the effect of domain shift on geometric statistics, leading to AUROC gains of +34 points in multi-domain settings.

幻觉检测几何信号大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。