arXiv:2605.10002cs.CV2026-05中稿 · IJCAI被引 2

首个3D肿瘤PET/CT多步推理幻觉检测基准,揭示医疗视觉模型系统性错误。

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

论文配图:Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
图 1 · 摘自论文原文
  • 分四阶段设计诊断流程,实现多步推理的细粒度幻觉检测
  • 覆盖12000+图像与百万级图文对,发现模型在复杂推理中幻觉率超60%
  • 适合医疗AI安全研究者,助力开发更可信的医学视觉模型

大型视觉语言模型(VLMs)在医学图像理解中表现强劲,但常生成看似合理却错误的临床陈述,引发严重安全问题。现有医疗幻觉基准主要聚焦2D影像的一次性诊断问题,难以揭示预测是否基于正确的定位与异常识别,导致关键推理错误被掩盖。我们提出Med-StepBench,首个面向3D肿瘤PET/CT的分步幻觉检测大规模基准,涵盖超过12,000张图像和超过1,000,000个图像-陈述对,涵盖体积分层与多视图2D数据,将临床推理分解为四个专家设计的诊断阶段。基于临床医生验证的标注,我们首次实现对通用及医学VLM的步骤级评估,揭示了被整体准确率掩盖的系统性失败模式。此外,我们发现当前VLM极易受到对抗性但临床上看似合理的中间解释影响,即使存在矛盾的视觉证据,仍会显著放大幻觉。研究结果突显了多步临床推理在模型泛化中的根本局限,并确立Med-StepBench作为开发更安全、更可靠的医疗VLM的严格基准。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet incorrect statements, raising significant safety concerns. Existing medical hallucination benchmarks primarily focus on 2D imaging with one-shot diagnostic questions, offering limited insight into whether predictions are grounded in correct localization and abnormality identification, allowing critical reasoning errors to remain hidden behind seemingly correct diagnoses. We introduce Med-StepBench, the first large-scale benchmark for step-wise hallucination detection in 3D oncological PET/CT, comprising over 12,000 images and more than 1,000,000 image-statement pairs across volumetric and multi-view 2D data, which decomposes clinical reasoning into four expert-designed diagnostic stages. Using clinician-verified annotations, we perform the first step-level evaluation of general-purpose and medical VLMs, revealing systematic failure modes obscured by aggregate accuracy metrics. Furthermore, we show that current VLMs are highly susceptible to adversarial yet clinically plausible intermediate explanations, which significantly amplify hallucinations despite contradictory visual evidence. Together, our findings highlight fundamental limitations in grounding multi-step clinical reasoning and establish Med-StepBench as a rigorous benchmark for developing safer and more reliable medical VLMs.

医疗AI幻觉检测多模态3D影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。