构建多时间点胸片纵向推理基准,检验模型对疾病演变的跨时序理解能力。
MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays

- 设计五次随访的多时段胸片问答任务,覆盖时间事件定位等三类临床推理场景。
- 14个先进视觉语言模型平均准确率仅29.3%,远低于随机水平,暴露其时序建模缺陷。
- 适用于评估医疗AI在长期病程分析中的表现,尤其适合研究纵向视觉推理的团队。
纵向胸片(CXR)解读需对多次就诊中疾病演变进行推理,但现有医学VQA基准多聚焦单图或短时序图像对。本文提出MI-CXR,一个用于标准化评估多时段纵向推理的基准,无需自由文本报告或额外临床信息。MI-CXR包含针对五次随访患者时间线的五选一多项选择题,涵盖三类互补任务:时间事件定位、区间变化推理与全局轨迹总结,均聚焦于临床相关的时序视觉推理。对14个前沿视觉语言模型(VLMs)的评估显示,整体表现较低,平均准确率为29.3%,仅略高于随机猜测。通过分阶段诊断探查发现,模型常生成局部合理的区间描述,却无法施加时间约束,或整合证据形成全时间线一致的判断。这些发现揭示了当前VLM在纵向推理中的关键局限,并确立了MI-CXR作为纵向医疗推理的规范性基准。基准已公开于https://github.com/AIDASLab/MI-CXR。
原文摘要 · Abstract (English)
Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA benchmarks focus on single images or short-horizon image pairs. We introduce MI-CXR, a benchmark for standardized evaluation of Multi-Interval longitudinal reasoning over multi-visit CXR sequences, without requiring free-form report generation or additional clinical context. MI-CXR comprises five-way multiple-choice questions over five-visit patient timelines and instantiates three complementary task families: Temporal Event Localization, Interval-wise Change Reasoning, and Global Trajectory Summarization, which assess clinically grounded visual reasoning over time. Evaluating 14 state-of-the-art vision-language models (VLMs) shows low overall performance, with an average accuracy of 29.3%, only modestly above random guessing. Using stage-wise diagnostic probing, we find that models often produce locally plausible interval descriptions but fail to enforce temporal constraints or compose evidence into globally consistent decisions over the full timeline. These findings reveal key limitations of current VLMs and establish MI-CXR as a principled benchmark for longitudinal medical reasoning. The benchmark is available at https://github.com/AIDASLab/MI-CXR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。