arXiv:2503.02891cs.CVcs.AR2025-03中稿 · Neurocomputing, El…综述被引 74

系统梳理视觉Transformer在边缘设备上的压缩与加速方法。

Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies

  • 按模型压缩、软件推理工具、硬件加速三类梳理技术体系
  • 对比不同策略在精度、效率与硬件适配性间的权衡
  • 适合关注边缘部署的AI工程师与研究者参考

近年来,视觉变压器(ViTs)在图像分类、目标检测和分割等计算机视觉任务中展现出强大潜力。与依赖分层特征提取的卷积神经网络(CNNs)不同,ViTs将图像视为补丁序列,并利用自注意力机制。然而,其高计算复杂度和内存需求给资源受限的边缘设备部署带来挑战。为此,大量研究聚焦于模型压缩技术和面向硬件的加速策略。但目前仍缺乏对这些技术及其在精度、效率与硬件适应性之间权衡的系统性综述。本文填补这一空白,系统分析模型压缩方法、边缘推理软件工具及硬件加速策略,探讨其对性能的影响,指出关键挑战与新兴研究方向,推动ViT在图形处理器(GPUs)、专用集成电路(ASICs)和现场可编程门阵列(FPGAs)等边缘平台上的部署,旨在为高效边缘部署提供当代指南。

原文摘要 · Abstract (English)

In recent years, vision transformers (ViTs) have emerged as powerful and promising techniques for computer vision tasks such as image classification, object detection, and segmentation. Unlike convolutional neural networks (CNNs), which rely on hierarchical feature extraction, ViTs treat images as sequences of patches and leverage self-attention mechanisms. However, their high computational complexity and memory demands pose significant challenges for deployment on resource-constrained edge devices. To address these limitations, extensive research has focused on model compression techniques and hardware-aware acceleration strategies. Nonetheless, a comprehensive review that systematically categorizes these techniques and their trade-offs in accuracy, efficiency, and hardware adaptability for edge deployment remains lacking. This survey bridges this gap by providing a structured analysis of model compression techniques, software tools for inference on edge, and hardware acceleration strategies for ViTs. We discuss their impact on accuracy, efficiency, and hardware adaptability, highlighting key challenges and emerging research directions to advance ViT deployment on edge platforms, including graphics processing units (GPUs), application-specific integrated circuit (ASICs), and field-programmable gate arrays (FPGAs). The goal is to inspire further research with a contemporary guide on optimizing ViTs for efficient deployment on edge devices.

视觉Transformer边缘计算模型压缩硬件加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。