让医学影像对比问答模型更懂空间差异,提升病灶识别准确率
Location-Aware Pretraining for Medical Difference Visual Question Answering
- 引入位置感知预训练,通过指代表达等任务学习空间对齐的视觉表征
- 在胸部X光片上实现当前最佳性能,准确识别临床相关变化
- 适合医学影像分析、AI辅助诊断方向的研究者参考
差异性医学视觉问答模型需比较多张图像以识别具有临床意义的变化,依赖视觉编码器捕捉反映放射科医生对比诊断流程的细微视觉差异。然而,使用标准对比或分类目标训练的视觉编码器常无法捕获区分疾病进展与成像相关变异所需的微小变化。为此,我们提出一种位置感知预训练框架,融合自动指代表达(AREF)、基于场景的描述(GCAP)和条件自动指代表达(CAREF)。这些任务促进细粒度、空间定位的视觉表示学习。当与语言模型结合时,该方法在医学差异性视觉问答任务中达到当前最优表现,能准确识别并推理胸部X光片中的临床相关变化。
原文摘要 · Abstract (English)
Differential medical VQA models compare multiple images to identify clinically meaningful changes and rely on vision encoders to capture fine-grained visual differences that reflect radiologists' comparative diagnostic workflows. However, vision encoders trained using standard contrastive or classification objectives often fail to capture the subtle variations needed to distinguish true disease progression from acquisition-related variability. To address this limitation, we introduce a location-aware pretraining framework that incorporates automatic referring expressions (AREF), grounded captioning (GCAP), and conditional automatic referring expressions (CAREF). These tasks promote the learning of fine-grained, spatially grounded visual representations. When integrated with a language model, our approach achieves state-of-the-art performance on medical difference VQA by accurately identifying and reasoning about clinically relevant changes in chest X-ray images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。