arXiv:2604.17439cs.CV2026-04

40篇论文梳理非Transformer视觉模型,揭示高效替代方案。

Attention Is not Everything: Efficient Alternatives for Vision

论文配图:Attention Is not Everything: Efficient Alternatives for Vision
图 1 · 摘自论文原文
  • 按卷积、MLP、状态空间等分类,系统整理非Transformer方法
  • 对比效率、可扩展性、可解释性与鲁棒性,评估性能表现
  • 为未来视觉研究提供新方向,适合关注模型轻量化的读者

近年来计算机视觉的发展主要得益于基于Transformer的模型。然而,许多非Transformer方法依然表现优异,与Transformer模型形成直接竞争。本文对40篇相关论文进行综述,构建了全面的分类体系,将这些方法归类为基于卷积、MLP、状态空间等类别。从效率、可扩展性、可理解性及鲁棒性等多个维度进行分析,旨在呈现非Transformer方法的全貌,揭示其面临的挑战与未来机遇,为后续计算机视觉研究提供参考。

原文摘要 · Abstract (English)

Recently computer vision has seen advancements mainly thanks to Transformer-based models. However many non-Transformer methods are still doing well being a direct competition of Transformer-based models. This review tries to present a comprehensive taxonomy of such methods and organize these methods into categories like convolution-based models, MLP-based models, state-space-based and more. These methods are looked at in terms of how efficient they are, how well they scale, how easy they are to understand and how robust they are. A total of 40 papers were chosen for this study. The goal is to give a view of non-Transformer methods and find out what challenges and opportunities exist for future computer vision research.

非Transformer视觉模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。