arXiv:2509.25143cs.CVcs.CL2025-09被引 1

构建时序医学图像推理基准,评估模型跟踪病情变化能力

TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models

  • 设计多任务时序医疗视觉语言推理评测集
  • 多数大模型在时序推理中表现仅相当于随机猜测
  • 多模态检索增强可显著提升模型性能

现有医学推理基准多基于单次就诊的图像分析,偏离真实临床实践。本文提出 TEMMED-BENCH,一个面向时序医疗图像推理的多任务评测基准,包含三类任务:视觉问答(VQA)、报告生成与图像对选择,并附带超过17,000条实例的知识语料库。我们评估了12个大型视觉语言模型(含6个开源与6个专有模型),结果表明大多数模型无法有效分析患者跨时间的病情变化,在闭卷设置下大量表现接近随机猜测。通过引入医学领域内检索到的视觉与文本模态作为输入增强,我们发现多模态检索增强相比无检索或仅文本检索,在多数模型上带来显著提升,其中VQA任务平均提升2.59%。该基准反映当前大模型在时序医学图像推理中的局限性,同时指出多模态检索增强是值得探索的改进方向。

原文摘要 · Abstract (English)

Existing medical reasoning benchmarks for vision-language models primarily focus on analyzing a patient's condition based on an image from a single visit. However, this setting deviates significantly from real-world clinical practice, where doctors typically refer to a patient's historical conditions to provide a comprehensive assessment by tracking their changes over time. In this paper, we introduce TEMMED-BENCH, a multi-task benchmark designed for analyzing changes in patients' conditions between different clinical visits, which challenges large vision-language models (LVLMs) to reason over temporal medical images. TEMMED-BENCH consists of a test set comprising three tasks - visual question-answering (VQA), report generation, and image-pair selection - and a supplementary knowledge corpus of over 17,000 instances. With TEMMED-BENCH, we conduct an evaluation of twelve LVLMs, comprising six proprietary and six open-source models. Our results show that most LVLMs lack the ability to analyze patients' condition changes over temporal medical images, and a large proportion perform only at a random-guessing level in the closed-book setting. To enhance the tracking of condition changes, we explore augmenting the input with both retrieved visual and textual modalities in the medical domain. We also show that multi-modal retrieval augmentation yields notably higher performance gains than no retrieval and textual retrieval alone across most models on our benchmark, with the VQA task showing an average improvement of 2.59%. Overall, we compose a benchmark grounded on real-world clinical practice, and it reveals LVLMs' limitations in temporal medical image reasoning, as well as highlighting the use of multi-modal retrieval augmentation as a potentially promising direction worth exploring to address this challenge.

时序推理医疗AI视觉语言模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。