通过对话中不确定性突变检测幻觉,跨领域效果更优
Beyond In-Domain Detection: SpikeScore for Cross-Domain Hallucination Detection
- 基于多轮对话的不确定性突变设计新评分指标SpikeScore
- 在多个LLM和基准上实现跨领域幻觉检测性能领先
- 适合需要可靠幻觉检测的跨领域实际应用
幻觉检测对大语言模型在真实场景中的部署至关重要。现有方法在训练与测试数据同域时表现良好,但在跨域场景下泛化能力差。本文研究了一项被忽视的重要问题——可泛化的幻觉检测(GHD),即在单一领域数据上训练检测器,仍能保持在多种相关领域中的鲁棒性能。我们模拟了大模型初始回复后的多轮对话,发现幻觉引发的对话在不同领域中均表现出比事实性对话更大的不确定性波动。基于此现象,提出新评分指标SpikeScore,用于量化多轮对话中的突变幅度。理论分析与实证验证表明,SpikeScore在跨域场景下能有效区分幻觉与非幻觉响应。在多个大语言模型和基准上的实验显示,基于SpikeScore的检测方法优于代表性基线,并超越先进泛化方法,验证了其在跨域幻觉检测中的有效性。
原文摘要 · Abstract (English)
Hallucination detection is critical for deploying large language models (LLMs) in real-world applications. Existing hallucination detection methods achieve strong performance when the training and test data come from the same domain, but they suffer from poor cross-domain generalization. In this paper, we study an important yet overlooked problem, termed generalizable hallucination detection (GHD), which aims to train hallucination detectors on data from a single domain while ensuring robust performance across diverse related domains. In studying GHD, we simulate multi-turn dialogues following LLMs' initial response and observe an interesting phenomenon: hallucination-initiated multi-turn dialogues universally exhibit larger uncertainty fluctuations than factual ones across different domains. Based on the phenomenon, we propose a new score SpikeScore, which quantifies abrupt fluctuations in multi-turn dialogues. Through both theoretical analysis and empirical validation, we demonstrate that SpikeScore achieves strong cross-domain separability between hallucinated and non-hallucinated responses. Experiments across multiple LLMs and benchmarks demonstrate that the SpikeScore-based detection method outperforms representative baselines in cross-domain generalization and surpasses advanced generalization-oriented methods, verifying the effectiveness of our method in cross-domain hallucination detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。