用少量参数让通用分割模型高效适配医学影像,大幅降低标注成本。
SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation
- 引入可动态适应模态与任务的专家混合机制,保持原模型接口和推理效率。
- 仅用不到1%数据量的精选数据集训练,性能超越现有方法。
- 自动提示分词实现零标注依赖,适合临床快速部署的医学图像分割场景。
医学图像分割对临床诊断与定量分析至关重要,但受成像模态异质性和像素级标注成本高影响,仍具挑战。尽管通用交互分割模型如SAM取得显著进展,其在医学影像中的迁移仍面临两大瓶颈:(i) 缺乏针对模态与解剖结构特异性任务的自适应机制,限制了分布外场景的泛化能力;(ii) 当前医学适配方法在大规模异构数据集上全量微调,导致噪声监督、成本升高及负迁移。为此,我们提出SegMoTE,一种高效且自适应的医学图像分割框架。SegMoTE保留SAM的原始提示接口、高效推理与零样本泛化能力,仅引入少量可学习参数,实现跨模态与任务的动态适应。同时设计渐进式提示分词机制,实现完全自动分割,显著降低标注依赖。在小于现有大规模数据集1%的精选数据集MedSeg-HQ上训练,SegMoTE在多种成像模态与解剖任务中达到当前最优性能。它是首个在极低标注成本下实现高效、稳健、可扩展的通用分割模型医学适配方案,推动基础视觉模型在临床应用中的实际落地。
原文摘要 · Abstract (English)
Medical image segmentation is vital for clinical diagnosis and quantitative analysis, yet remains challenging due to the heterogeneity of imaging modalities and the high cost of pixel-level annotations. Although general interactive segmentation models like SAM have achieved remarkable progress, their transfer to medical imaging still faces two key bottlenecks: (i) the lack of adaptive mechanisms for modality- and anatomy-specific tasks, which limits generalization in out-of-distribution medical scenarios; and (ii) current medical adaptation methods fine-tune on large, heterogeneous datasets without selection, leading to noisy supervision, higher cost, and negative transfer. To address these issues, we propose SegMoTE, an efficient and adaptive framework for medical image segmentation. SegMoTE preserves SAM's original prompt interface, efficient inference, and zero-shot generalization while introducing only a small number of learnable parameters to dynamically adapt across modalities and tasks. In addition, we design a progressive prompt tokenization mechanism that enables fully automatic segmentation, significantly reducing annotation dependence. Trained on MedSeg-HQ, a curated dataset less than 1% of existing large-scale datasets, SegMoTE achieves SOTA performance across diverse imaging modalities and anatomical tasks. It represents the first efficient, robust, and scalable adaptation of general segmentation models to the medical domain under extremely low annotation cost, advancing the practical deployment of foundation vision models in clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。