arXiv:2502.01812cs.CLcs.LG2025-02被引 5

提出无需资源的多模块幻觉检测框架,专治大模型数学推理错误

SelfCheck-Eval: A Multi-Module Framework for Zero-Resource Hallucination Detection in Large Language Models

  • 构建三模块黑盒检测框架,融合语义、专业和上下文一致性分析
  • 在数学推理任务上现有方法准确率不足50%,暴露出检测短板
  • 首个专门针对数学幻觉的AIME数据集,适合高精度领域研究者使用

大语言模型在问答、医疗、法律等领域展现强大能力,但其生成虚构内容(幻觉)的问题严重制约了在高风险场景中的可靠应用。现有检测基准多聚焦通用知识,忽视了对数学推理等关键领域的评估。为此,我们构建了首个面向数学推理幻觉的AIME Math Hallucination数据集,并提出SelfCheck-Eval框架——一种不依赖模型内部结构、适用于开源与闭源模型的多模块检测系统。该框架整合语义、专业领域与上下文一致性三个独立检测模块。实验表明,当前方法在人物传记类内容中表现良好,但在数学推理任务上性能显著下降,即便采用NLI微调、偏好学习和过程监督等主流策略,准确率仍低于50%。结果揭示现有检测技术在数学领域的根本性局限,凸显开发专用、黑盒兼容方法的紧迫性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse applications, from open-domain question answering to scientific writing, medical decision support, and legal analysis. However, their tendency to generate incorrect or fabricated content, commonly known as hallucinations, represents a critical barrier to reliable deployment in high-stakes domains. Current hallucination detection benchmarks are limited in scope, focusing primarily on general-knowledge domains while neglecting specialised fields where accuracy is paramount. To address this gap, we introduce the AIME Math Hallucination dataset, the first comprehensive benchmark specifically designed for evaluating mathematical reasoning hallucinations. Additionally, we propose SelfCheck-Eval, a LLM-agnostic, black-box hallucination detection framework applicable to both open and closed-source LLMs. Our approach leverages a novel multi-module architecture that integrates three independent detection strategies: the Semantic module, the Specialised Detection module, and the Contextual Consistency module. Our evaluation reveals systematic performance disparities across domains: existing methods perform well on biographical content but struggle significantly with mathematical reasoning, a challenge that persists across NLI fine-tuning, preference learning, and process supervision approaches. These findings highlight the fundamental limitations of current detection methods in mathematical domains and underscore the critical need for specialised, black-box compatible approaches to ensure reliable LLM deployment.

幻觉检测数学推理大模型安全黑盒评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。