arXiv:2505.14726eess.IVcs.AI2025-05被引 10

微调BLIP模型提升医学影像描述准确性

MedBLIP: Fine-tuning BLIP for Medical Image Captioning

  • 在ROCO数据集上微调BLIP,增强医学图像理解能力
  • 全模型微调效果最佳,生成描述准确率显著提升
  • 仅微调解码器可节省5%训练时间,适合资源受限场景

医学影像描述是一项挑战性任务,要求生成临床准确且语义清晰的放射科图像描述。尽管近期视觉语言模型(如BLIP、BLIP2、Gemini和ViT-GPT2)在自然图像数据集上表现优异,但在医学领域常生成通用或不精确的描述。本研究探索在ROCO数据集上微调BLIP模型以提升放射科图像描述性能。对比了微调后的BLIP与零样本版本、BLIP-2 base、BLIP-2 Instruct及ViT-GPT2基线模型。结果表明,针对医学领域的微调显著提升了量化与定性评估指标表现。通过可视化解码器交叉注意力图评估可解释性,并开展消融实验分析编码器与解码器单独微调的贡献。研究强调了医学应用中针对性适配的重要性,发现仅微调解码器(编码器冻结)可实现5%更短训练时间,提供良好性能基线,而全模型微调仍为最优方案。

原文摘要 · Abstract (English)

Medical image captioning is a challenging task that requires generating clinically accurate and semantically meaningful descriptions of radiology images. While recent vision-language models (VLMs) such as BLIP, BLIP2, Gemini and ViT-GPT2 show strong performance on natural image datasets, they often produce generic or imprecise captions when applied to specialized medical domains. In this project, we explore the effectiveness of fine-tuning the BLIP model on the ROCO dataset for improved radiology captioning. We compare the fine-tuned BLIP against its zero-shot version, BLIP-2 base, BLIP-2 Instruct and a ViT-GPT2 transformer baseline. Our results demonstrate that domain-specific fine-tuning on BLIP significantly improves performance across both quantitative and qualitative evaluation metrics. We also visualize decoder cross-attention maps to assess interpretability and conduct an ablation study to evaluate the contributions of encoder-only and decoder-only fine-tuning. Our findings highlight the importance of targeted adaptation for medical applications and suggest that decoder-only fine-tuning (encoder-frozen) offers a strong performance baseline with 5% lower training time than full fine-tuning, while full model fine-tuning still yields the best results overall.

医学影像图像描述BLIP微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。