提出新方法量化模型检测幻觉时对问题信息的依赖,发现现有方法多靠‘作弊’而非真认知。
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
- 提出无须人工标注的近似问题侧效应(AQE)评估方法
- 发现现有幻觉检测模型性能主要依赖于题干信息而非真实理解
- 适合关注模型真实能力评估的研究者与开发者参考
许多现有语言模型幻觉检测方法报告了看似强劲的性能表现。然而我们指出,这些性能不仅反映模型对其内部知识的真实认知,还包含仅来自问题侧信息(如基准测试作弊)的虚假信号。尽管此类作弊在现有基准上可提升检测分数,但无法泛化到域外场景或实际应用。然而,区分模型性能中来自问题侧信号的比例非常困难。为此,我们提出一种无需人工标注的量化方法——近似问题侧效应(AQE)。基于AQE的分析显示,当前主流幻觉检测方法严重依赖基准作弊,而非真正的自我意识。
原文摘要 · Abstract (English)
Many works have proposed methodologies for language model (LM) hallucination detection and reported seemingly strong performance. However, we argue that the reported performance to date reflects not only a model's genuine awareness of its internal information, but also awareness derived purely from question-side information (e.g., benchmark hacking). While benchmark hacking can be effective for boosting hallucination detection score on existing benchmarks, it does not generalize to out-of-domain settings and practical usage. Nevertheless, disentangling how much of a model's hallucination detection performance arises from question-side awareness is non-trivial. To address this, we propose a methodology for measuring this effect without requiring human labor, Approximate Question-side Effect (AQE). Our analysis using AQE reveals that existing hallucination detection methods rely heavily on benchmark hacking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。