arXiv:2606.21517cs.CL2026-06

医学幻觉检测器的定位能力需实证检验,而非仅依赖架构宣称。

MedHal-Loc: Are "Explainable-by-Architecture" Medical Hallucination Detectors Faithful Localizers? A Localization Benchmark

论文配图:MedHal-Loc: Are "Explainable-by-Architecture" Medical Hallucination Detectors Faithful Localizers? A Localization Benchmark
图 1 · 摘自论文原文
  • 构建新基准MedHal-Loc,量化检测器定位错误片段的能力。
  • 59%实体抽取覆盖率导致知识图谱方法定位效果仅略高于随机(+3.3pp)。
  • 真实幻觉多为模糊结论反转,难定位,适合关注可解释性的研究者。

临床文本幻觉检测正被视作可解释性问题:系统不仅需标记不可靠内容,还需指出具体错误片段。基于知识图谱三元组分解的架构被宣传具备审计能力,但其定位性能常被假设而非验证。本文提出MedHal-Loc基准与评估指标,衡量检测器高分错误单元是否真实覆盖错误片段。控制集包含300个来自PubMedQA的陈述,单片段错误按四类可定位类型(实体替换、关系错误、机制误归因、虚构)注入,金标准片段由构造决定;自然集表明真实幻觉以广泛结论反转为主,人类专家仅接受1/18候选片段。评估四种细粒度范式发现:逐句NLI、逐句一致性及专用跨度检测器FAVA定位显著优于随机,而复杂知识图谱三元组流水线定位仅略高于随机(+3.3个百分点,不显著),受限于约59%的实体抽取覆盖率——尽管检测F1达0.609。检测能力不等于可靠定位,架构可解释性须实证验证。

原文摘要 · Abstract (English)

Detecting hallucinations in clinical text is increasingly framed as an explainability problem: systems should not merely flag an unreliable response but point to the offending span. Architectures built around knowledge-graph (KG) triple decomposition are marketed for exactly this auditability, yet their localization ability is typically assumed rather than measured. We introduce MedHal-Loc, a benchmark and metric for localization faithfulness -- whether a detector's top-ranked error unit actually overlaps the erroneous span. The controlled subset comprises 300 PubMedQA-derived statements with single, span-level errors injected across four localizable types (entity substitution, relation error, mechanism misattribution, invention), yielding gold spans by construction; a complementary natural subset documents that real hallucinations are dominated by diffuse conclusion-flips that resist span localization (a human expert accepted 1/18 candidate spans). Evaluating four fine-grained paradigms, we find that NLI-per-clause, consistency-per-sentence, and the dedicated span detector FAVA all localize well above chance, whereas an elaborate KG-triple pipeline localizes no better than chance (+3.3pp, n.s.), bottlenecked by ~59% entity-extraction coverage -- despite competitive detection F1 (0.609). Detection competence does not imply faithful localization; architectural explainability must be validated, not presumed.

医学幻觉可解释性定位评估知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。