提出SeVeR框架,高效处理多序列脑部MRI的冗余视觉信息。
SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

- 按模态压缩体积数据,仅保留关键原型
- 用感知变化的门控注意力检索互补证据,减少85%以上视觉标记
- 适合医学影像问答中需精准推理的场景
三维医学图像问答面临长且冗余的视觉标记序列问题,尤其在多序列磁共振成像中,不同模态提供互补诊断线索却导致解码器重复接触相同解剖区域。为此,我们首先构建了BreMRIs-VQA,一个临床筛选的乳腺MRI基准数据集,包含119万组问答对、7.1万条序列和1.29万名患者,涵盖自由文本与多选题。进一步提出SeVeR框架,通过模态专属原型压缩密集体积,并在解码时利用变化感知的门控注意力检索多层次互补证据,采用边际效用自一致性目标训练,抑制无益检索。在BreMRIs-VQA及公开基准上的实验表明,SeVeR显著提升判别与生成性能,同时大幅减少视觉标记暴露量。
原文摘要 · Abstract (English)
Volumetric medical VQA requires reasoning over long and redundant 3D visual token sequences, especially in multi-sequence MRI where complementary modalities provide diverse diagnostic cues but expose the decoder to many repeated anatomical regions. To investigate reasoning under multi-sequence visual redundancy, we first introduce BreMRIs-VQA, a clinically curated breast MRI benchmark with 1.19M QA pairs from 71.0K sequences and 12.9K patients, covering both free-text and multiple-choice questions. We further propose SeVeR, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval. Experiments on BreMRIs-VQA and public benchmarks show that SeVeR improves both discriminative and generative performance while exposing substantially fewer visual tokens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。