arXiv:2505.16964cs.CVcs.CL2025-05被引 31

首个测试多图医学问答的基准,专为临床推理设计。

MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning

  • 用医学教育视频脚本自动对齐图像与文本,生成多图问答对。
  • 11个先进模型在多图推理上准确率普遍低于50%,表现不稳定。
  • 适合评估下一代医学多模态模型的时序推理能力。

真实临床实践需要多图对比推理,但现有医学基准仍局限于单帧理解。我们提出MedFrameQA,首个专为多图医学视觉问答(VQA)设计的基准,基于教育验证的诊断序列构建。通过开发可扩展的流水线,利用医学教育视频的叙事脚本将视觉帧与文本概念对齐,自动生成2,851组高质量的多图VQA对,包含明确、基于脚本的推理链。对11个先进多模态大模型(包括推理模型)的评估显示,其在多图融合方面存在严重缺陷,准确率大多低于50%,且在不同图像数量下表现不稳定。错误分析表明,模型常将图像视为孤立实例,无法追踪病灶演变或跨图像进行解剖参照。MedFrameQA为评估下一代医学多模态大模型处理复杂、时序关联的医学叙述提供了严格标准。

原文摘要 · Abstract (English)

Real-world clinical practice demands multi-image comparative reasoning, yet current medical benchmarks remain limited to single-frame interpretation. We present MedFrameQA, the first benchmark explicitly designed to test multi-image medical VQA through educationally-validated diagnostic sequences. To construct this dataset, we develop a scalable pipeline that leverages narrative transcripts from medical education videos to align visual frames with textual concepts, automatically producing 2,851 high-quality multi-image VQA pairs with explicit, transcript-grounded reasoning chains. Our evaluation of 11 advanced MLLMs (including reasoning models) exposes severe deficiencies in multi-image synthesis, where accuracies mostly fall below 50% and exhibit instability across varying image counts. Error analysis demonstrates that models often treat images as isolated instances, failing to track pathological progression or cross-reference anatomical shifts. MedFrameQA provides a rigorous standard for evaluating the next generation of MLLMs in handling complex, temporally grounded medical narratives.

医学视觉问答多图推理临床决策多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。