用合成数据微调MedGemma,让医学图像描述更准确可信。
Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
- 用知识蒸馏生成合成数据,通过QLoRA高效微调模型。
- 分类准确率提升,且描述忠实度、相关性显著改善。
- 适合医疗RAG系统开发者,提升临床决策支持能力。
检索增强生成系统在提供马来西亚临床实践指南的基于事实的指导中至关重要,但其对图像查询的效果有限,因通用视觉-语言模型的描述常缺乏临床特异性与事实依据。本研究提出并验证了一个框架,将MedGemma模型专门化以生成高保真度的图像描述,作为更优查询。为应对数据稀缺问题,采用知识蒸馏流程,在皮肤科、眼底和胸部放射领域构建合成数据集,并使用参数高效的QLoRA方法微调模型。性能通过双框架评估:一是分类准确率,二是创新应用RAGAS框架评估描述的忠实度、相关性与正确性。微调后模型在分类表现上显著提升,且RAGAS评估显示描述忠实度与正确性大幅提高,验证了模型生成可靠、事实依据充分描述的能力。该工作建立了一个可靠的医学视觉-语言模型专业化流程,并验证其作为高质量查询生成器的有效性,为提升循证临床决策支持中的多模态RAG系统奠定基础。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation systems are essential for providing fact-based guidance from Malaysian Clinical Practice Guidelines. However, their effectiveness with image-based queries is limited, as general Vision-Language Model captions often lack clinical specificity and factual grounding. This study proposes and validates a framework to specialize the MedGemma model for generating high-fidelity captions that serve as superior queries. To overcome data scarcity, we employ a knowledge distillation pipeline to create a synthetic dataset across dermatology, fundus, and chest radiography domains, and fine-tune MedGemma using the parameter-efficient QLoRA method. Performance was rigorously assessed through a dual framework measuring both classification accuracy and, via a novel application of the RAGAS framework, caption faithfulness, relevancy, and correctness. The fine-tuned model demonstrated substantial improvements in classification performance, while RAGAS evaluation confirmed significant gains in caption faithfulness and correctness, validating the models ability to produce reliable, factually grounded descriptions. This work establishes a robust pipeline for specializing medical VLMs and validates the resulting model as a high-quality query generator, laying the groundwork for enhancing multimodal RAG systems in evidence-based clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。