用对比学习提升胎盘分析模型效率,更省资源更准确
VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
- 用文本锚定的对比知识蒸馏优化医疗视觉语言模型
- 压缩后模型性能不降反升,加速明显且适配低质图像
- 适合医疗边缘设备部署,尤其资源受限场景
胎盘病理检查是发现和降低分娩相关健康风险的有效手段。近年来,人工智能利用胎盘照片和病理报告实现了分娩相关病变的检测与分类。然而,现有自动化方法计算开销大,限制了实际部署。本文提出对视觉-语言对比学习(VLC)框架的两项改进:(1) 文本锚定的视觉-语言对比知识蒸馏(VLCD),一种用于医学VLC预训练的新知识蒸馏策略;(2) 使用大规模自然图像数据集进行无监督预蒸馏,以改善模型初始化。该方法蒸馏出高效神经网络,在性能上达到或超过教师模型,同时实现模型压缩与加速。结果表明,无监督预蒸馏显著提升了模型在低质量图像下的表现与鲁棒性。VLCD有效提升了医学VLC方法的效率与可部署性,使AI医疗解决方案在资源受限环境中更具可及性。
原文摘要 · Abstract (English)
Pathological examination of the placenta is an effective method for detecting and mitigating health risks associated with childbirth. Recent advancements in AI have enabled the use of photographs of the placenta and pathology reports for detecting and classifying signs of childbirth-related pathologies. However, existing automated methods are computationally extensive, which limits their deployability. We propose two modifications to vision-language contrastive learning (VLC) frameworks to enhance their accuracy and efficiency: (1) text-anchored vision-language contrastive knowledge distillation (VLCD)-a new knowledge distillation strategy for medical VLC pretraining, and (2) unsupervised predistillation using a large natural images dataset for improved initialization. Our approach distills efficient neural networks that match or surpass the teacher model in performance while achieving model compression and acceleration. Our results showcase the value of unsupervised predistillation in improving the performance and robustness of our approach, specifically for lower-quality images. VLCD serves as an effective way to improve the efficiency and deployability of medical VLC approaches, making AI-based healthcare solutions more accessible, especially in resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。