arXiv:2604.15808cs.CVcs.AI2026-04

构建多帧医学MRI空间定位问答数据集,推动视觉语言模型具备临床级推理能力。

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI

论文配图:Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
图 1 · 摘自论文原文
  • 基于专家标注的4万+对多帧MRI问答,含帧索引框坐标链式推理
  • 引入10个视觉语言模型对比,显示边界框监督显著提升定位精度
  • 适合医学影像分析、可解释性AI研究者关注临床级多帧推理

空间推理与视觉定位是视觉语言模型的核心能力,但多数医学VLM缺乏透明推理过程或空间证据。现有评估基准仅基于孤立2D图像,忽略了临床影像的三维特性——病灶可能跨多个切片或仅出现在少数层面。本文提出面向体积化MRI的空间定位问答基准SGMRI-VQA,包含41,307对来自fastMRI+数据集脑部和膝关节研究的专家放射科医师标注,每对问答均配有与临床一致的思维链及帧索引的边界框坐标。任务按检测、定位、计数/分类、描述四层组织,要求模型共同推断内容、位置及其跨帧分布。我们对10个VLM进行评测,发现对Qwen3-VL-8B进行有监督微调并加入边界框监督,能持续优于强零样本基线,表明针对性的空间监督是实现临床可解释推理的有效路径。

原文摘要 · Abstract (English)

Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without transparent reasoning or spatial evidence. Existing benchmarks also evaluate VLMs on isolated 2D images, overlooking the volumetric nature of clinical imaging, where findings can span multiple frames or appear on only a few slices. We introduce Spatially Grounded MRI Visual Question Answering (SGMRI-VQA), a 41,307-pair benchmark for multi-frame, spatially grounded reasoning on volumetric MRI. Built from expert radiologist annotations in the fastMRI+ dataset across brain and knee studies, each QA pair includes a clinician-aligned chain-of-thought trace with frame-indexed bounding box coordinates. Tasks are organized hierarchically across detection, localization, counting/classification, and captioning, requiring models to jointly reason about what is present, where it is, and across which frames it extends. We benchmark 10 VLMs and show that supervised fine-tuning of Qwen3-VL-8B with bounding box supervision consistently improves grounding performance over strong zero-shot baselines, indicating that targeted spatial supervision is an effective path toward grounded clinical reasoning.

医学影像空间推理多帧理解可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。