用视觉基础模型提升显微图像像素与物体分类效果
Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy
- 结合浅层学习与注意力探测,评估多种视觉基础模型
- 在5个数据集上均优于传统手工特征方法
- 为显微图像分类提供新基准,适合医学图像研究者
深度学习支撑了现代计算机视觉的多数方法,包括生物医学成像。然而,在交互式语义分割(即像素分类)和交互式物体级分类中,基于特征的浅层学习仍广泛使用,原因在于该领域数据多样性高、缺乏大规模预训练数据集,且对计算与标注效率要求高。相比之下,显微镜领域许多任务(如细胞实例分割)已采用深度学习,并因视觉基础模型(VFMs)特别是SAM的引入而显著受益。本文探究了VFMs能否在像素与物体分类任务中超越现有方法。我们评估了包括通用模型(SAM、SAM2、SAM3、DINOv3)和领域特定模型(μSAM、PathoSAM、KRONOS)在内的多种模型,结合浅层学习与注意力探测,在五个多样且具有挑战性的数据集上进行测试。结果表明,相比手工特征方法,各类模型均实现持续改进,并为实际应用提供了清晰路径。此外,本研究建立了显微镜领域视觉基础模型的基准,推动未来该方向的发展。
原文摘要 · Abstract (English)
Deep learning underlies most modern approaches and tools in computer vision, including biomedical imaging. However, for interactive semantic segmentation (often called pixel classification in this context) and interactive object-level classification (object classification), feature-based shallow learning remains widely used. This is due to the diversity of data in this domain, the lack of large pretraining datasets, and the need for computational and label efficiency. In contrast, state-of-the-art tools for many other vision tasks in microscopy - most notably cellular instance segmentation - already rely on deep learning and have recently benefited substantially from vision foundation models (VFMs), particularly SAM. Here, we investigate whether VFMs can also improve pixel and object classification compared to current approaches. To this end, we evaluate several VFMs, including general-purpose models (SAM, SAM2, SAM3, DINOv3) and domain-specific ones ($μ$SAM, PathoSAM, KRONOS), in combination with shallow learning and attentive probing on five diverse and challenging datasets. Our results demonstrate consistent improvements over hand-crafted features and provide a clear pathway toward practical improvements. Furthermore, our study establishes a benchmark for VFMs in microscopy and informs future developments in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。