用统一框架自动分割多种眼底图像,精度媲美专家模型。
CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging
- 用CLIP预训练图像编码器,融合多模态眼底数据提升泛化能力。
- 通过GroundingDINO自动生成病灶框提示,结合模态标识文本实现精准分割。
- 全流程自动化,适配11类任务、12个数据集,无需人工调参。
近年来,基础模型如分割一切模型(SAM)在医学图像分割中取得显著进展,尤其在眼底成像领域,精确分割对诊断至关重要。然而现有方法仍面临三大挑战:1)文本疾病描述存在模态歧义;2)依赖人工生成提示;3)缺乏统一框架,多数方法局限于特定模态与任务。为此,我们提出CLIP统一自动提示分割(CLAPS),一种面向眼底成像多任务、多模态的统一分割方法。首先,在大规模多模态眼底数据集上预训练基于CLIP的图像编码器,缓解数据稀缺与分布不均问题。随后,利用GroundingDINO自动检测局部病灶并生成空间边界框提示。为统一任务并消除歧义,采用带“模态标识符”的文本提示。最终,自动化的文本与空间提示引导SAM完成高精度分割,构建端到端自动化统一流程。在12个不同数据集、11类关键分割任务上的大量实验表明,CLAPS性能与专用专家模型相当,且在多数指标上超越现有基准,验证了其作为基础模型的广泛泛化能力。
原文摘要 · Abstract (English)
Recent advancements in foundation models, such as the Segment Anything Model (SAM), have significantly impacted medical image segmentation, especially in retinal imaging, where precise segmentation is vital for diagnosis. Despite this progress, current methods face critical challenges: 1) modality ambiguity in textual disease descriptions, 2) a continued reliance on manual prompting for SAM-based workflows, and 3) a lack of a unified framework, with most methods being modality- and task-specific. To overcome these hurdles, we propose CLIP-unified Auto-Prompt Segmentation (\CLAPS), a novel method for unified segmentation across diverse tasks and modalities in retinal imaging. Our approach begins by pre-training a CLIP-based image encoder on a large, multi-modal retinal dataset to handle data scarcity and distribution imbalance. We then leverage GroundingDINO to automatically generate spatial bounding box prompts by detecting local lesions. To unify tasks and resolve ambiguity, we use text prompts enhanced with a unique "modality signature" for each imaging modality. Ultimately, these automated textual and spatial prompts guide SAM to execute precise segmentation, creating a fully automated and unified pipeline. Extensive experiments on 12 diverse datasets across 11 critical segmentation categories show that CLAPS achieves performance on par with specialized expert models while surpassing existing benchmarks across most metrics, demonstrating its broad generalizability as a foundation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。