arXiv:2502.14584eess.IVcs.CV2025-02被引 11

大模型赋能医学图像分割,突破数据少与领域差异瓶颈

Vision Foundation Models in Medical Image Analysis: Advances and Challenges

  • 用适配器与知识蒸馏改进视觉基础模型,提升医学图像适应性
  • 结合多尺度特征建模,显著提升小样本下的分割精度
  • 适合关注医学影像智能分析的科研与临床人员阅读

视觉基础模型(VFMs),特别是视觉变压器(ViT)和通用分割模型(SAM),在医学图像分析领域取得显著进展,展现出捕捉长程依赖和高泛化能力。然而,将这些大模型应用于医学图像面临诸多挑战:医学图像与自然图像存在领域差异、需高效适配策略,且医学数据集规模有限。本文综述了当前视觉基础模型在医学图像分割中的最新研究进展,重点聚焦于领域自适应、模型压缩与联邦学习。文章讨论了基于适配器的改进、知识蒸馏技术及多尺度上下文特征建模等方法,并提出未来发展方向。分析表明,结合联邦学习与模型压缩等新兴方法,有望推动医学图像分析变革,增强临床应用价值。本文旨在系统梳理现有技术路径,为下一阶段创新提供关键研究方向。

原文摘要 · Abstract (English)

The rapid development of Vision Foundation Models (VFMs), particularly Vision Transformers (ViT) and Segment Anything Model (SAM), has sparked significant advances in the field of medical image analysis. These models have demonstrated exceptional capabilities in capturing long-range dependencies and achieving high generalization in segmentation tasks. However, adapting these large models to medical image analysis presents several challenges, including domain differences between medical and natural images, the need for efficient model adaptation strategies, and the limitations of small-scale medical datasets. This paper reviews the state-of-the-art research on the adaptation of VFMs to medical image segmentation, focusing on the challenges of domain adaptation, model compression, and federated learning. We discuss the latest developments in adapter-based improvements, knowledge distillation techniques, and multi-scale contextual feature modeling, and propose future directions to overcome these bottlenecks. Our analysis highlights the potential of VFMs, along with emerging methodologies such as federated learning and model compression, to revolutionize medical image analysis and enhance clinical applications. The goal of this work is to provide a comprehensive overview of current approaches and suggest key areas for future research that can drive the next wave of innovation in medical image segmentation.

医学图像视觉模型分割适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。