用基础模型适配器实现跨模态医学图像少样本多标签分割
Data Adaptive Few-shot Multi Label Segmentation with Foundation Model
- 基于视觉变换器的适配器,利用子像素特征匹配目标区域
- 在2D/3D及不同体位临床数据上均优于现有少样本分割方法
- 支持单图模板处理多标签任务,适用于真实医疗场景
图像分割与定位的精准标注成本高昂,促使少样本算法成为研究热点。尽管现有基于文本提示的方法在部分任务中表现优异,但在医学图像上仍存在性能瓶颈。利用基于视觉变换器(ViT)的基础模型的子像素级特征,可通过单张模板图像有效实现跨模态医学图像的一次性分割与定位。然而,此类方法依赖模板与测试图像高度匹配,且仅靠简单相关性难以建立准确对应关系。在实际临床数据中,因患者体位变化、同一模态内扫描协议差异或三维数据扩展,该方法泛化能力受限。此外,多标签任务需逐个识别感兴趣区域(RoI),效率较低。本文提出一种基于基础模型(FM)的适配器框架,可同时支持单标签与多标签的定位与分割。实验表明,该方法在2D与3D数据、不同体位临床数据上均显著优于当前最先进的少样本分割方法。
原文摘要 · Abstract (English)
The high cost of obtaining accurate annotations for image segmentation and localization makes the use of one and few shot algorithms attractive. Several state-of-the-art methods for few-shot segmentation have emerged, including text-based prompting for the task but suffer from sub-optimal performance for medical images. Leveraging sub-pixel level features of existing Vision Transformer (ViT) based foundation models for identifying similar region of interest (RoI) based on a single template image have been shown to be very effective for one shot segmentation and localization in medical images across modalities. However, such methods rely on assumption that template image and test image are well matched and simple correlation is sufficient to obtain correspondences. In practice, however such an approach can fail to generalize in clinical data due to patient pose changes, inter-protocol variations even within a single modality or extend to 3D data using single template image. Moreover, for multi-label tasks, the RoI identification has to be performed sequentially. In this work, we propose foundation model (FM) based adapters for single label, multi-label localization and segmentation to address these concerns. We demonstrate the efficacy of the proposed method for multiple segmentation and localization tasks for both 2D and 3D data as we well as clinical data with different poses and evaluate against the state of the art few shot segmentation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。