arXiv:2605.24792cs.CVcs.AI2026-05

用轻量微调技术提升胃肠道内镜AI的问答与隐私保护图像生成能力。

Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering

论文配图:Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
图 1 · 摘自论文原文
  • 采用低秩适配(LoRA)实现模型轻量化微调,降低90%计算成本。
  • 合成图像达到0.290保真度、0.730一致性,弗雷歇生物医学CLIP距离仅1450。
  • 适合医疗AI研究者和临床工程师,解决数据稀缺与隐私难题。

胃肠道内镜AI系统受限于标注数据不足、隐私政策严格及传统微调效率低下,难以在临床落地。本文提出双管道参数高效微调(PEFT)框架,解决临床视觉问答(VQA)与隐私保护合成数据生成两大问题。基于Florence-2模型,结合低秩适配(LoRA)显著提升可解释性并降低训练开销;同时利用Stable Diffusion 2.1与LoRA生成高质量胃肠道图像,增强训练集且不泄露患者隐私。实验基于Kvasir-VQA数据集,所提VQA模型在ROUGE-1达0.92、ROUGE-L达0.91,BLEU从0.08提升至0.24。在私有数据集上微调效果优于公共数据集。秩为4的LoRA合成图像在保真度(0.290)、一致率(0.730)和弗雷歇生物医学CLIP距离(FBD=1450)方面表现最优,计算成本减少近90%。相较FLUX、MSDM与Kandinsky 2.2,本方法在图像-文本一致性上更优,验证了其在临床AI中的可行性与鲁棒性。

原文摘要 · Abstract (English)

The major limitations of gastrointestinal (GI) endoscopy AI systems arise from a shortage of annotated data, strict privacy policies, and significant bottlenecks in conventional model fine-tuning. Such limitations impede the successful application of sophisticated AI models in clinical practice, particularly affecting the reliability and scalability of diagnosis. In this paper, we present a dual-pipeline PEFT model that addresses two fundamental problems: medical Visual Question Answering (VQA) and the generation of privacy-preserving synthetic data. For clinical VQA, we adopt the Florence-2 vision-language model. Leveraging PEFT enhances model interpretability while substantially reducing the computational cost of training. Simultaneously, we employ Low-Rank Adaptation (LoRA) with Stable Diffusion 2.1 to generate high-quality GI images that enhance training databases without violating patient privacy. This research utilized the Kvasir-VQA dataset. Our Florence-2 VQA model achieved ROUGE-1 of 0.92, ROUGE-L of 0.91, and BLEU score improvements from 0.08 to 0.24. Fine-tuning on private datasets consistently showed better results than fine-tuning on public datasets. The rank-4 LoRA synthesis achieved optimal performance with a fidelity score of 0.290, an agreement score of 0.730, and a Frechet BiomedCLIP Distance (FBD) of 1450, reducing computational costs by almost 90 percent. This framework improves the clinical potential of AI in GI endoscopy. Compared to FLUX, MSDM, and Kandinsky 2.2, our model demonstrates superior FBD and strong semantic alignment. While other models lead in Fidelity or Agreement, our lower FBD indicates better image-text coherence. These results establish our approach as a robust solution for enhancing VQA and synthetic data generation in clinical AI.

医疗AI轻量微调图像生成视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。