arXiv:2501.09372cs.CV2025-01综述被引 2

Transformer取代CNN,解决图像分割中的长距离依赖和尺度变化问题。

Image Segmentation with transformers: An Overview, Challenges and Future

  • 用Transformer捕捉图像中远距离像素的依赖关系
  • 克服CNN在多尺度物体分割上的局限性
  • 适合研究视觉分割与模型轻量化方向的读者

图像分割是计算机视觉的核心任务,传统上依赖卷积神经网络(CNN),但这类模型难以捕捉复杂的空间依赖性、处理不同尺度的目标、需要人工设计网络结构,并缺乏上下文信息建模能力。本文分析了基于CNN模型的不足,探讨向Transformer架构的转变以克服这些限制。综述了当前先进的基于Transformer的分割模型,讨论了分割任务特有的挑战及其解决方案。同时指出当前面临的难题,并展望未来趋势,如轻量级架构与数据效率提升。本综述为理解Transformer如何推动分割技术发展、突破传统模型局限提供了指引。

原文摘要 · Abstract (English)

Image segmentation, a key task in computer vision, has traditionally relied on convolutional neural networks (CNNs), yet these models struggle with capturing complex spatial dependencies, objects with varying scales, need for manually crafted architecture components and contextual information. This paper explores the shortcomings of CNN-based models and the shift towards transformer architectures -to overcome those limitations. This work reviews state-of-the-art transformer-based segmentation models, addressing segmentation-specific challenges and their solutions. The paper discusses current challenges in transformer-based segmentation and outlines promising future trends, such as lightweight architectures and enhanced data efficiency. This survey serves as a guide for understanding the impact of transformers in advancing segmentation capabilities and overcoming the limitations of traditional models.

图像分割Transformer深度学习视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。