arXiv:2510.05189cs.CLcs.AI2025-10

用提示工程生成幻觉数据,通过向量空间分析实现轻量级幻觉检测。

A novel hallucination classification framework

  • 基于提示工程系统生成多种幻觉类型,构建专用幻觉数据集。
  • 幻觉在向量空间中与正确答案集群距离越远,信息失真越严重。
  • 仅用简单分类算法即可可靠区分幻觉与真实输出,适合快速部署。

本文提出一种自动检测大语言模型推理过程中生成幻觉的新方法。该方法基于系统性分类框架,通过提示工程可控地复现多种幻觉类型,并将构建的幻觉数据集映射到嵌入向量空间。利用降维后的向量表示,结合无监督学习技术分析幻觉与真实响应的分布特征。定量评估显示,幻觉与正确输出聚类中心之间的距离与其信息失真程度呈稳定正相关。这一发现为即使简单的分类算法也能在单个大语言模型中可靠区分幻觉与真实回答提供了理论与实证支持,从而构建了一个轻量但高效的幻觉检测框架。

原文摘要 · Abstract (English)

This work introduces a novel methodology for the automatic detection of hallucinations generated during large language model (LLM) inference. The proposed approach is based on a systematic taxonomy and controlled reproduction of diverse hallucination types through prompt engineering. A dedicated hallucination dataset is subsequently mapped into a vector space using an embedding model and analyzed with unsupervised learning techniques in a reduced-dimensional representation of hallucinations with veridical responses. Quantitative evaluation of inter-centroid distances reveals a consistent correlation between the severity of informational distortion in hallucinations and their spatial divergence from the cluster of correct outputs. These findings provide theoretical and empirical evidence that even simple classification algorithms can reliably distinguish hallucinations from accurate responses within a single LLM, thereby offering a lightweight yet effective framework for improving model reliability.

幻觉检测大语言模型向量空间无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。