arXiv:2609.03953cs.CL2026-09

多视角判别能更全面发现医疗聊天机器人幻觉,但结果仍依赖专家判断。

Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

  • 设计多阶段标注流程,融合专家与证据核查
  • 单次标注漏检率高,大模型辅助发现更多候选错误
  • 现有基准仍会遗漏错误,需多源判别提升覆盖

理解聊天机器人生成文本中事实性错误的频率并评估检测系统至关重要,关乎聊天机器人安全性。然而,事实性错误检测常被视为单次、单标注任务。在长篇回复中,错误可能隐晦且嵌于多数正确内容之中。本文开展医学相关聊天机器人回复的多视角标注研究,结合初筛标注、大模型作为裁判(LaJ)候选发现,以及医学专家与基于证据的事实核查两种裁定方式。初筛标注员常遗漏后续裁定者验证的错误;LaJ虽提升候选发现效率,但无法独立胜任——它会漏掉标注员捕捉到的错误。同时,裁定者间存在分歧,表明多源候选判定可增强基准完整性,但仍需依赖判断与专业能力。应用于现有基准时,该方法同样揭示了标注遗漏现象。结果表明,在本研究设定下,单次标注基准虽具规模,却可能低估事实错误;多轮裁定可提高覆盖率,但最终结论仍受判断标准、专业知识和证据支持影响。

原文摘要 · Abstract (English)

Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-checking. First-pass annotators frequently miss factual errors later validated by adjudicators. LaJ improves candidate discovery, but is insufficient on its own: It misses factual errors that annotators catch. We also find disagreement among adjudicators, suggesting that adjudication over multiple candidate sources can improve benchmark completeness, but does not eliminate the need to apply judgment and expertise. Applied to an existing benchmark, this technique reveals a similar pattern of missing annotations. Together, these results suggest that in the settings examined here, single-pass hallucination benchmarks may achieve scale at the cost of undercounting factual errors. Multi-pass adjudication can improve coverage, but inferences drawn from the benchmarks are still sensitive to the judgment, expertise, and evidence used to determine error presence.

医疗AI幻觉检测多视角判别标注质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。