arXiv:2505.21786cs.CLcs.AI2025-05被引 3

让AI生成内容的错误可追溯,提升可信度。

VeriTrail: Closed-Domain Hallucination Detection with Traceability

  • 通过追踪每一步生成内容,定位虚假信息源头。
  • 在多步生成任务中检测准确率优于现有方法。
  • 首个支持全过程可追溯的幻觉检测工具,适合AI安全研究者。

即使被要求严格遵循源材料,语言模型仍常生成无根据的内容,这种现象称为“封闭域幻觉”。在多步生成过程(MGS)中,该风险比单步生成(SGS)更高。然而,由于MGS过程更复杂,仅检测最终输出中的幻觉已不够:还需追踪幻觉内容可能在何处引入,以及内容如何从源材料经中间步骤忠实生成。为此,我们提出VeriTrail,首个专为MGS和SGS设计的封闭域幻觉检测方法,并构建首个包含所有中间输出及人工标注最终输出忠实性的数据集。实验证明,VeriTrail在两个数据集上均优于基线方法。

原文摘要 · Abstract (English)

Even when instructed to adhere to source material, language models often generate unsubstantiated content - a phenomenon known as "closed-domain hallucination." This risk is amplified in processes with multiple generative steps (MGS), compared to processes with a single generative step (SGS). However, due to the greater complexity of MGS processes, we argue that detecting hallucinations in their final outputs is necessary but not sufficient: it is equally important to trace where hallucinated content was likely introduced and how faithful content may have been derived from the source material through intermediate outputs. To address this need, we present VeriTrail, the first closed-domain hallucination detection method designed to provide traceability for both MGS and SGS processes. We also introduce the first datasets to include all intermediate outputs as well as human annotations of final outputs' faithfulness for their respective MGS processes. We demonstrate that VeriTrail outperforms baseline methods on both datasets.

幻觉检测可追溯性多步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。