arXiv:2605.02443cs.CL2026-05

构建系统化评测框架,精准检测和缓解大模型指令跟随中的幻觉问题。

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs

  • 设计多维度评测框架,覆盖6种检测方法与3个领域
  • 提出新指标HalluScore,与人工评估相关性达0.41
  • 开发自适应路由算法,成本减半仅损失0.1%性能

大型语言模型在自然语言处理任务中表现卓越,但仍易产生幻觉——生成与事实不符、脱离上下文或违背用户指令的内容。我们提出HalluScan,一个全面的基准测试框架,系统评估72种配置下幻觉检测与缓解效果,涵盖6种检测方法、4个开源模型家族和3个不同领域。主要贡献包括:(1) HalluScore,一种新型综合指标,与人工专家判断的皮尔逊相关系数达r=0.41;(2) 自适应检测路由(ADR),智能路由算法实现2.0倍成本降低,仅导致0.1% AUROC下降;(3) 系统性错误级联分解,揭示不同领域间幻觉类型存在显著差异。实验表明,NLI验证方法取得最高总体AUROC为0.88,RAV次之,为0.66。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks, yet they remain susceptible to hallucinations -- generating content that is factually incorrect, unfaithful to provided context, or misaligned with user instructions. We present HalluScan, a comprehensive benchmark framework that systematically evaluates hallucination detection and mitigation across 72 configurations spanning 6 detection methods, 4 open-weight model families, and 3 diverse domains. We introduce three key contributions: (1) HalluScore, a novel composite metric that achieves a Pearson correlation of r = 0.41 with human expert judgments; (2) Adaptive Detection Routing (ADR), an intelligent routing algorithm achieving 2.0x cost reduction with only 0.1% AUROC degradation; and (3) systematic error cascade decomposition revealing substantial variation in hallucination error types across domains. Our experiments reveal that NLI Verification achieves the highest overall AUROC of 0.88, while RAV achieves the second-highest AUROC of 0.66.

幻觉检测大模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。