低资源下医学图像分类新框架,提升小样本与零样本表现
Multi-View Synergistic Learning with Vision-Language Adaption for Low-Resource Biomedical Image Classification

- 分步适配视觉与语言编码器,实现高效参数微调
- 多粒度对比学习增强病变区域判别力,准确率提升显著
- 利用大模型结构化监督,保持疾病语义一致性,适合医疗场景
在标注数据稀缺的低资源条件下,精准进行医学图像分类仍具挑战,主要源于标注有限、类间视觉差异细微及疾病语义复杂。尽管视觉-语言模型可缓解数据不足问题,但其在医学场景中的有效适配受限于参数高效微调与细粒度、语义一致表征学习的需求。本文提出多视图协同学习(MVSL)框架,通过联合考虑适配范式、表征粒度与疾病语义关系,解决上述问题。MVSL解耦视觉与文本编码器的适配过程,尊重其不同表征特性,实现更稳定高效的参数高效微调;引入多粒度对比学习,显式建模全局图像语义与局部病灶级证据,提升对视觉相似疾病的细粒度区分能力;同时,通过大语言模型生成的结构化监督,保留疾病级别的语义结构,约束文本表征并间接正则化视觉嵌入。实验在11个公开医学数据集(涵盖9种成像模态、10个解剖区域)上验证,MVSL在少样本与零样本分类设置中持续优于现有最优方法。
原文摘要 · Abstract (English)
Accurate biomedical image classification under low-resource conditions remains challenging due to limited annotations, subtle inter-class visual differences, and complex disease semantics. While vision--language models offer a promising foundation for mitigating data scarcity, their effective adaptation in biomedical settings is constrained by the need for parameter-efficient tuning alongside fine-grained and semantically consistent representation learning. In this work, we propose Multi-View Synergistic Learning (MVSL), a unified framework that addresses these challenges by jointly considering adaptation paradigms, representation granularity, and disease semantic relationships. MVSL decouples the adaptation of visual and textual encoders to respect their distinct representational characteristics, enabling more stable and effective parameter-efficient fine-tuning. It further introduces multi-granularity contrastive learning to explicitly model both global image semantics and localized lesion-level evidence, improving fine-grained discrimination for visually similar disease categories. In addition, MVSL preserves disease-level semantic structure by incorporating structured supervision derived from large language models, which constrains textual representations at the class level and indirectly regularizes visual embeddings through cross-modal alignment. Together, these components enable more stable cross-modal alignment and improved discrimination under limited supervision. Extensive experiments on $11$ public biomedical datasets spanning $9$ imaging modalities and $10$ anatomical regions demonstrate that MVSL consistently outperforms state-of-the-art methods in few-shot and zero-shot classification settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。