构建统一基准评估少样本医学图像分割,揭示模型优劣关键
Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation

- 设计统一基准FAME,覆盖多种主流方法与多模态数据
- 14,958个测试样本验证:直接视觉适配优于提示策略
- 适用于医疗影像领域研究者,尤其关注少样本泛化能力
少样本医学图像分割(FS-MIS)旨在仅凭少量标注的支撑样本分割新的感兴趣区域。尽管进展迅速,现有方法范式多样且评估标准不一,导致其相对效果不清晰。本文提出FAME,一个统一的FS-MIS评估基准,涵盖专家模型、SAM-based方法、CLIP-based方法和MLLM-based方法。FAME包含跨7个解剖部位、9种成像模态、14类目标区域的14,958个测试样本,评估零样本与十样本设置,并额外考察目标不存在识别能力及在协变量与语义偏移下的泛化性能。实验发现:有效分割依赖于模型对支撑样本的利用方式,直接视觉适应通常优于基于提示的策略;增加支撑样本仅在模型能有效利用时才提升性能;语义迁移仍远难于成像域适应,强定位能力不等于可靠的无目标识别。希望FAME为当前FS-MIS方法提供全面理解,并推动更高效可靠的少样本医学分割发展。
原文摘要 · Abstract (English)
Few-shot medical image segmentation (FS-MIS) aims to segment novel regions of interest (ROIs) from a few annotated support examples. Despite rapid progress, existing FS-MIS solutions span diverse paradigms but are evaluated under inconsistent settings, leaving their relative effectiveness unclear. We introduce FAME, a unified benchmark for evaluating FS-MIS solutions, covering specialists, SAM-based methods, CLIP-based methods, and MLLM-based methods. FAME contains 14,958 test samples across 7 anatomical sites, 9 imaging modalities, and 14 ROI categories, and evaluates models under zero-shot and ten-shot settings with additional assessment of target-absence recognition and generalization under covariate and semantic shifts. Our evaluation reveals several findings. First, effective few-shot segmentation depends on how models exploit support examples: direct visual adaptation generally outperforms prompt-based strategies. Second, increasing support examples improves performance only when models can effectively utilize them. Third, semantic transfer remains substantially more challenging than imaging-domain adaptation, and strong localization ability does not necessarily imply reliable target-absence recognition. We hope FAME provides a comprehensive understanding of current FS-MIS solutions and facilitates the development of more effective and reliable few-shot medical segmentation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。