arXiv:2509.13813cs.CL2025-09被引 12

用几何方法量化大模型回答的不确定性,更好识别幻觉。

Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs

  • 基于响应嵌入的原型分析,构建全局与局部不确定性度量。
  • 在医学数据集上优于现有方法,显著降低幻觉率。
  • 无需模型内部信息,适用于任意黑箱大模型。

大型语言模型在多项任务中表现优异,但仍会生成看似合理却错误的回答(即幻觉)。不确定性量化被视为检测幻觉的有效策略,需同时估计全局不确定性(整批回答)和局部不确定性(单个回答)。现有黑箱方法多依赖零散启发式或图论近似,缺乏统一几何解释。本文提出一种几何框架,仅通过黑箱访问采样响应批次,进行原型分析。全局层面引入几何体积(Geometric Volume),衡量由响应嵌入导出的原型凸包体积;局部层面提出几何怀疑度(Geometric Suspicion),利用响应与原型的空间关系对可靠性排序,支持通过优选回答实现幻觉抑制。相比依赖离散成对比较的方法,本方法提供连续语义边界点,更精细地分配个体可靠性。实验表明,该框架在短问答数据集上表现相当或更优,在医疗数据集上显著优于现有方法,因幻觉风险更高而更具价值。理论证明了凸包体积与熵之间的关联。

原文摘要 · Abstract (English)

Large language models demonstrate impressive results across diverse tasks but are still known to hallucinate, generating linguistically plausible but incorrect answers to questions. Uncertainty quantification has been proposed as a strategy for hallucination detection, requiring estimates for both global uncertainty (attributed to a batch of responses) and local uncertainty (attributed to individual responses). While recent black-box approaches have shown some success, they often rely on disjoint heuristics or graph-theoretic approximations that lack a unified geometric interpretation. We introduce a geometric framework to address this, based on archetypal analysis of batches of responses sampled with only black-box model access. At the global level, we propose Geometric Volume, which measures the convex hull volume of archetypes derived from response embeddings. At the local level, we propose Geometric Suspicion, which leverages the spatial relationship between responses and these archetypes to rank reliability, enabling hallucination reduction through preferential response selection. Unlike prior methods that rely on discrete pairwise comparisons, our approach provides continuous semantic boundary points which have utility for attributing reliability to individual responses. Experiments show that our framework performs comparably to or better than prior methods on short form question-answering datasets, and achieves superior results on medical datasets where hallucinations carry particularly critical risks. We also provide theoretical justification by proving a link between convex hull volume and entropy.

幻觉检测不确定性几何方法大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。