arXiv:2607.27122cs.CV2026-07中稿 · EMA4MICCAI 2026

让小模型更懂内窥镜图像,回答时能说出依据

Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

论文配图:Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
图 1 · 摘自论文原文
  • 用已有标注数据构造定位和描述任务,少需额外标注
  • 小模型在多任务下准确率提升,且答案更贴近图像区域
  • 适合需要可解释性的医疗AI应用

胃肠道(GI)内窥镜图像分析正从单标签分类转向视觉问答(VQA),要求模型能回答关于图像的自由形式临床问题。尽管现有视觉语言模型(VLM)在该任务上表现良好,但临床应用还需模型的答案有视觉证据支撑。本文提出一种简单的多任务微调方法:利用已有数据集中的专家标注息肉掩膜作为直接监督,对无真值掩膜的类别则使用基于Grad-CAM的弱监督进行定位。在Kvasir-VQA-x1数据集上,三种小型VLM骨干网络采用低秩适配(LoRA)进行微调,对比仅任务与多任务训练方案,结果表明多任务设置在分布内和分布外数据上均实现稳定准确率提升,并增强了答案词元与相关图像区域之间的隐式对齐。

原文摘要 · Abstract (English)

Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must answer free-form clinical questions about an image. While recent vision-language models (VLMs) achieve promising answer accuracy on this task, clinical adoption also requires the model's internal representations to reflect the visual evidence behind its answers. We propose a simple multi-task fine-tuning recipe that constructs auxiliary grounding and description tasks from an existing VQA dataset with minimal additional annotation: expert-annotated polyp masks are reused directly, while a GI-domain pretrained classifier with Grad-CAM localization provides weak supervision for finding categories that lack ground-truth masks. Three small VLM backbones are fine-tuned with low-rank adaptation under matched VQA-only and multi-task recipes on Kvasir-VQA-x1, and we show consistent accuracy gains together with improved implicit alignment between answer tokens and the relevant image region, evaluated on both in-distribution and out-of-distribution data.

医疗视觉视觉问答可解释性小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。