用分层对比学习让大模型高效分析少样本医学影像
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
- 分两阶段微调,通过层次化对齐视觉与文本特征
- 在胸片和乳腺超声数据集上达到顶尖少样本与零样本效果
- 适合医疗图像分析、小样本学习研究者参考
医学图像分类中的少样本学习因标注数据稀缺和图像复杂性而面临巨大挑战。本文提出自适应视觉-语言微调与分层对比对齐(HiCA)框架,利用大视觉-语言模型(LVLMs)进行医学图像分析。该方法采用两阶段微调策略,结合领域特定预训练与分层对比学习,在多层级上对齐视觉与文本表征。我们在胸片和乳腺超声两个基准数据集上评估了该方法,在少样本与零样本设置下均达到当前最优性能。进一步分析表明,该方法具有强鲁棒性、良好泛化能力与可解释性,相比现有基线有显著提升。本工作凸显了分层对比策略在适配LVLM应对医学影像任务独特挑战方面的潜力。
原文摘要 · Abstract (English)
Few-shot learning in medical image classification presents a significant challenge due to the limited availability of annotated data and the complex nature of medical imagery. In this work, we propose Adaptive Vision-Language Fine-tuning with Hierarchical Contrastive Alignment (HiCA), a novel framework that leverages the capabilities of Large Vision-Language Models (LVLMs) for medical image analysis. HiCA introduces a two-stage fine-tuning strategy, combining domain-specific pretraining and hierarchical contrastive learning to align visual and textual representations at multiple levels. We evaluate our approach on two benchmark datasets, Chest X-ray and Breast Ultrasound, achieving state-of-the-art performance in both few-shot and zero-shot settings. Further analyses demonstrate the robustness, generalizability, and interpretability of our method, with substantial improvements in performance compared to existing baselines. Our work highlights the potential of hierarchical contrastive strategies in adapting LVLMs to the unique challenges of medical imaging tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。