改进Transformer模型,提升显微图像语义分割精度。
Going Beyond U-Net: Assessing Vision Transformers for Semantic Segmentation in Microscopy Image Analysis

- 通过架构优化解决Swin Transformer在显微图像中的局限性
- 改进后模型性能超越U-Net和原始Swin-UPerNet
- 适合医学图像分析、生物成像领域的研究者参考
语义分割是显微图像分析的关键步骤。尽管U-Net仍是生物医学分割任务中最流行且成熟的方法,但基于Transformer的模型近年来展现出提升分割效果的潜力。本文评估了UNETR、Segment Anything Model及Swin-UPerNet等Transformer模型,并与经典U-Net在电子显微镜、明场、组织病理学和相位对比等多种图像模态下进行对比。研究发现原始Swin Transformer存在若干局限性,通过架构改进显著提升了其性能。实验结果表明,优化后的模型在多个数据集上优于传统U-Net和未修改的Swin-UPerNet。该对比分析证实,经过精心调整的Transformer模型在生物医学图像分割中具有巨大潜力,可推动其在显微图像分析工具中的应用。
原文摘要 · Abstract (English)
Segmentation is a crucial step in microscopy image analysis. Numerous approaches have been developed over the past years, ranging from classical segmentation algorithms to advanced deep learning models. While U-Net remains one of the most popular and well-established models for biomedical segmentation tasks, recently developed transformer-based models promise to enhance the segmentation process of microscopy images. In this work, we assess the efficacy of transformers, including UNETR, the Segment Anything Model, and Swin-UPerNet, and compare them with the well-established U-Net model across various image modalities such as electron microscopy, brightfield, histopathology, and phase-contrast. Our evaluation identifies several limitations in the original Swin Transformer model, which we address through architectural modifications to optimise its performance. The results demonstrate that these modifications improve segmentation performance compared to the classical U-Net model and the unmodified Swin-UPerNet. This comparative analysis highlights the promise of transformer models for advancing biomedical image segmentation. It demonstrates that their efficiency and applicability can be improved with careful modifications, facilitating their future use in microscopy image analysis tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。