arXiv:2504.02351cs.CVcs.AI2025-04被引 2

用大模型知识蒸馏,让小模型在医学图像分割上更准更快。

Agglomerating Large Vision Encoders via Distillation for VFSS Segmentation

  • 从多个专业大模型蒸馏知识,融合多任务优势
  • 12个分割任务平均提升2%的Dice系数
  • 适合资源有限但需高精度分割的研究者

基础模型在医学影像领域已取得显著成效,但其下游任务训练开销大、推理复杂度高。尽管已有轻量化版本,但受限于模型容量和训练策略,性能不佳。为此,我们提出一种新框架,通过从多个医学大模型(如MedSAM、RAD-DINO、MedCLIP)中蒸馏知识,这些模型各自擅长不同视觉任务,旨在有效缩小小模型在医学图像分割任务上的性能差距。聚合后的模型在12项分割任务上表现出更强泛化能力,而专用模型需为每项任务单独训练。相比简单蒸馏,该方法平均提升2%的Dice系数。

原文摘要 · Abstract (English)

The deployment of foundation models for medical imaging has demonstrated considerable success. However, their training overheads associated with downstream tasks remain substantial due to the size of the image encoders employed, and the inference complexity is also significantly high. Although lightweight variants have been obtained for these foundation models, their performance is constrained by their limited model capacity and suboptimal training strategies. In order to achieve an improved tradeoff between complexity and performance, we propose a new framework to improve the performance of low complexity models via knowledge distillation from multiple large medical foundation models (e.g., MedSAM, RAD-DINO, MedCLIP), each specializing in different vision tasks, with the goal to effectively bridge the performance gap for medical image segmentation tasks. The agglomerated model demonstrates superior generalization across 12 segmentation tasks, whereas specialized models require explicit training for each task. Our approach achieved an average performance gain of 2\% in Dice coefficient compared to simple distillation.

医学图像知识蒸馏模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。