arXiv:2605.23629cs.CV2026-05被引 1

评测AI医生诊断过程,而非只看最终答案。

DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs

论文配图:DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs
图 1 · 摘自论文原文
  • 让模型逐步请求影像数据,动态更新诊断概率
  • 211个病例中,多数模型未获取关键证据即下结论
  • 适合研究临床推理与多模态AI的学者

医学诊断不是从完整病历中一次做出判断,而是一个循序渐进的过程:医生决定收集哪些证据,不断修正鉴别诊断,并在诊断充分支持时停止。当前大多数医疗AI评测会提前暴露所有信息,仅评分最终答案,导致盲目猜测、过早闭合、低效检查和错误不确定性更新等问题被掩盖。我们提出DDX-TRACE,一个由医师评审的多模态神经影像学诊断轨迹基准,包含211个高难度病例,在证据隐藏条件下评估模型的诊断路径。每例从有限临床史开始,模型自由请求影像检查,获得匹配图像包后更新概率性鉴别诊断,再决定是否停止并给出定位诊断。评估主流视觉语言模型发现,最终诊断得分可能严重扭曲工作流程质量:模型可能在缺乏关键证据时猜测合理诊断,请求有用检查但误读原始图像,或低效获取证据且未能合理更新不确定性。通过控制性证据变体,可分离出规划、视觉证据提取和下游推理中的瓶颈。DDX-TRACE推动医疗AI评价从单一答案转向证据支撑的诊断轨迹。

原文摘要 · Abstract (English)

Medical diagnosis is not a single prediction from a fully specified vignette. It is a sequential workup: clinicians decide what evidence to obtain, revise a differential diagnosis, and stop when the diagnosis is sufficiently supported. Most medical AI benchmarks instead reveal the relevant context upfront and score only the final answer, making unsupported correct guesses, premature closure, inefficient workups, and poor uncertainty updating invisible. We introduce DDX-TRACE, a physician-adjudicated benchmark for multimodal neuroradiology that evaluates diagnostic trajectories under hidden evidence over 211 challenging cases. Each case begins with limited clinical history; models request imaging studies in free form, receive matched image bundles when available, update a probabilistic differential diagnosis after each turn, and stop with a localized final diagnosis. Evaluating state-of-the-art VLMs, we find that final diagnosis scores can substantially misrepresent workup quality: models may guess plausible diagnoses without essential evidence, request useful studies but misinterpret raw images, or acquire evidence inefficiently while updating uncertainty poorly. Controlled evidence variants isolate bottlenecks in planning, visual evidence extraction, and downstream differential reasoning. DDX-TRACE shifts medical AI evaluation from final answers to evidence-supported diagnostic trajectories.

医疗AI诊断轨迹多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。