用通用视觉模型的潜在规律,提升医学图像分割效果
LUMOS: Latent Universal Medical Priors for Segmentation
- 冻结通用视觉模型提取低层视觉先验,供医学分割使用
- 在多个医学数据集上实现显著性能提升,无需额外标注
- 适合资源有限的医学图像分割任务,尤其适合小样本场景
通用视觉基础模型(VFMs)主要基于自然图像训练,其在医学图像分割中的应用常被认为需昂贵的适配或领域特定微调。本文从新视角出发:并非让分割模型重新学习视觉规律,而是探究这些用于解剖分割的底层视觉先验是否已潜藏于通用模型中。我们发现,即使未接受医学监督,冻结的通用视觉模型仍编码可迁移的视觉规律,且这些规律不仅适用于自然图像,也适用于医学图像理解。为此,我们提出一种新框架LUMOS(Latent Universal Medical PriOrs for Segmentation),通过两个关键组件放大通用视觉模型的先验能力:(1) Pathfinder从冻结的视觉基础模型中提炼视觉线索;(2) Inspiror利用提炼出的视觉规律为传统医学网络提供空间引导。该方法使分割器无需从有限医学标注中完全学习复杂视觉规律,而能聚焦于任务相关的解剖轮廓识别。在多种医学数据集和基于标记的通用视觉模型上,实验表明当冻结的标记空间保留了块级模式相关性时,通用视觉模型可作为稳定的空间先验生成器。DINO在骨干结构匹配下表现稳定,而SigLIP则因标记粒度和表示目标差异展现出模型特异性敏感性。
原文摘要 · Abstract (English)
General vision foundation models (VFMs) have been primarily developed on natural images, and their utility for medical image segmentation is therefore often considered to depend on costly adaptation or domain-specific fine-tuning. In this paper, we revisit this assumption from a different perspective: rather than requiring VFM segmentors to relearn visual regularities, we investigate whether the low-level visual priors necessary for anatomical delineation already lie dormant within general VFMs. We observe that frozen VFMs, despite lacking medical supervision, encode transferable visual regularities. These properties are not exclusive to natural images but are also fundamental to medical image understanding. Motivated by this observation, we propose Latent Universal Medical PriOrs for Segmentation (LUMOS), a novel framework that amplifies general VFM priors to conventional medical segmentors. LUMOS consists of two key components: (1) Pathfinder that distills visual cues from a frozen vision foundation model, and (2) Inspiror that sparks the conventional medical networks with spatial guidance from distilled visual regularities. In this way, the segmentor is relieved from learning complex visual regularities entirely from limited medical annotations and can instead focus on task-specific anatomical delineation. Across diverse medical datasets and token-based VFMs, LUMOS shows that general VFMs can serve as spatial prior generators when their frozen token spaces preserve patch-level pattern relevance. DINO provides stable matched-backbone gains, while SigLIP exposes VFM-specific sensitivity caused by its different token granularity and representation objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。