无需参考文本,通过一致性与对齐模式检测大模型幻觉。
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
- 结合响应间与查询-响应一致性,提升检测能力。
- 在多领域数据集上平均检测置信度超90%。
- 适合需要可靠幻觉检测的闭源大模型应用。
基于响应一致性的无参考幻觉检测(RFHD)方法不依赖内部模型状态(如生成概率或梯度),适用于无法访问的闭源大语言模型。然而,其难以捕捉查询-响应对齐模式,常导致检测准确率较低。此外,缺乏覆盖多领域的大型基准数据集,现有数据集规模和范围有限。为此,我们提出HalluCounter,一种新型无参考幻觉检测方法,同时利用响应-响应与查询-响应的一致性及对齐模式,训练分类器以检测幻觉并提供置信度分数和最优响应。同时,我们构建了HalluCounterEval基准数据集,包含多个领域中合成与人工标注样本。所提方法显著优于现有最先进方法,在各数据集上平均检测置信度超过90%。
原文摘要 · Abstract (English)
Response consistency-based, reference-free hallucination detection (RFHD) methods do not depend on internal model states, such as generation probabilities or gradients, which Grey-box models typically rely on but are inaccessible in closed-source LLMs. However, their inability to capture query-response alignment patterns often results in lower detection accuracy. Additionally, the lack of large-scale benchmark datasets spanning diverse domains remains a challenge, as most existing datasets are limited in size and scope. To this end, we propose HalluCounter, a novel reference-free hallucination detection method that utilizes both response-response and query-response consistency and alignment patterns. This enables the training of a classifier that detects hallucinations and provides a confidence score and an optimal response for user queries. Furthermore, we introduce HalluCounterEval, a benchmark dataset comprising both synthetically generated and human-curated samples across multiple domains. Our method outperforms state-of-the-art approaches by a significant margin, achieving over 90\% average confidence in hallucination detection across datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。