arXiv:2601.09734cs.CLcs.AI2026-01AAAI被引 1

用自动数据生成提升大模型幻觉诊断能力,让错误可定位、可解释、可修复。

From Detection to Diagnosis: Advancing Hallucination Analysis with Automated Data Synthesis

  • 构建自动化数据合成流水线,生成带诊断标签的幻觉样本。
  • 训练出40亿参数模型,诊断性能媲美更大通用模型。
  • 首次实现幻觉的定位、归因与修正,适合可信AI研发者使用。

大语言模型的幻觉(即生成与事实或上下文不符的内容)是其在关键领域可靠部署的核心障碍。现有研究多聚焦于二元检测,虽能识别幻觉,却无法提供可解释、可操作的反馈,限制了实际应用。为此,本文提出从“检测”到“诊断”的新范式,引入幻觉诊断任务——要求模型不仅识别幻觉,还需定位错误、解释成因并修正内容。我们开发了幻觉诊断生成器(HDG),通过受控事实伪造和推理链扰动等多维增强策略,从原始语料中自动生成高质量带诊断元数据的训练样本。基于此数据,我们训练了HDM-4B-RL(40亿参数),采用包含结构、准确率和定位信号的综合奖励函数进行组相对策略优化(GRPO)。实验表明,该模型在HaluEval基准上超越此前最优检测模型,且在综合诊断任务中达到与更大通用模型相当的性能,同时保持更小规模。本工作验证了幻觉诊断的可行性与价值,为构建更可信的生成式AI系统提供了有效方法。

原文摘要 · Abstract (English)

Hallucinations in Large Language Models (LLMs), defined as the generation of content inconsistent with facts or context, represent a core obstacle to their reliable deployment in critical domains. Current research primarily focuses on binary "detection" approaches that, while capable of identifying hallucinations, fail to provide interpretable and actionable feedback for model improvement, thus limiting practical utility. To address this limitation, a new research paradigm is proposed, shifting from "detection" to "diagnosis". The Hallucination Diagnosis Task is introduced, a task which requires models to not only detect hallucinations, but also perform error localization, causal explanation, and content correction. We develop the Hallucination Diagnosis Generator (HDG), an automated pipeline that systematically generates high-quality training samples with rich diagnostic metadata from raw corpora through multi-dimensional augmentation strategies including controlled fact fabrication and reasoning chain perturbation. Using HDG-generated data, we train HDM-4B-RL, a 4-billion-parameter hallucination diagnosis model, employing Group Relative Policy Optimization (GRPO) with a comprehensive reward function incorporating structural, accuracy, and localization signals. Experimental results demonstrate that our model surpasses previous state-of-the-art detection models on the HaluEval benchmark while achieving comparable performance to advanced general-purpose models. In comprehensive diagnosis tasks, HDM-4B-RL matches the capabilities of larger general models while maintaining a smaller size. This work validates the feasibility and value of hallucination diagnosis, providing an effective methodology for building more trustworthy and reliable generative AI systems.

幻觉诊断数据合成大模型评估可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。