arXiv:2505.13584cs.CV2025-05综述被引 5

自监督学习让图像分割无需人工标注也能高效训练

Self-Supervised Learning for Image Segmentation: A Comprehensive Survey

  • 通过设计伪任务利用海量无标签图像学习有效特征
  • 覆盖150+篇论文,系统梳理自监督分割方法体系
  • 适合想进入计算机视觉领域的研究者快速入门

监督学习需要大量精确标注数据才能取得良好效果,而数据标注耗时耗力成本高。自监督学习(SSL)通过利用大量无标签数据并设计代理任务来学习有用表征,无需人工标注即可提升模型性能。该方法已成为解决图像分类、目标检测和分割等实际视觉任务的重要范式。图像分割是医学影像、智能交通、农业和安防等高级视觉应用的核心。尽管基于自监督的语义分割算法仍有巨大研究潜力,但系统性梳理现有方法对追踪进展和引导新研究者至关重要。本综述深入分析了超过150篇近期图像分割相关文献,重点聚焦自监督学习。它提出了代理任务、下游任务及常用基准数据集的实用分类体系,并从大量文献中提炼关键发现,展望未来方向,以使该领域更易理解与进入。

原文摘要 · Abstract (English)

Supervised learning demands large amounts of precisely annotated data to achieve promising results. Such data curation is labor-intensive and imposes significant overhead regarding time and costs. Self-supervised learning (SSL) partially overcomes these limitations by exploiting vast amounts of unlabeled data and creating surrogate (pretext or proxy) tasks to learn useful representations without manual labeling. As a result, SSL has become a powerful machine learning (ML) paradigm for solving several practical downstream computer vision problems, such as classification, detection, and segmentation. Image segmentation is the cornerstone of many high-level visual perception applications, including medical imaging, intelligent transportation, agriculture, and surveillance. Although there is substantial research potential for developing advanced algorithms for SSL-based semantic segmentation, a comprehensive study of existing methodologies is essential to trace advances and guide emerging researchers. This survey thoroughly investigates over 150 recent image segmentation articles, particularly focusing on SSL. It provides a practical categorization of pretext tasks, downstream tasks, and commonly used benchmark datasets for image segmentation research. It concludes with key observations distilled from a large body of literature and offers future directions to make this research field more accessible and comprehensible for readers.

自监督学习图像分割计算机视觉综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。