40篇论文梳理非Transformer视觉模型,揭示高效替代方案。
Attention Is not Everything: Efficient Alternatives for Vision

- 按卷积、MLP、状态空间等分类,系统整理非Transformer方法
- 对比效率、可扩展性、可解释性与鲁棒性,评估性能表现
- 为未来视觉研究提供新方向,适合关注模型轻量化的读者
近年来计算机视觉的发展主要得益于基于Transformer的模型。然而,许多非Transformer方法依然表现优异,与Transformer模型形成直接竞争。本文对40篇相关论文进行综述,构建了全面的分类体系,将这些方法归类为基于卷积、MLP、状态空间等类别。从效率、可扩展性、可理解性及鲁棒性等多个维度进行分析,旨在呈现非Transformer方法的全貌,揭示其面临的挑战与未来机遇,为后续计算机视觉研究提供参考。
原文摘要 · Abstract (English)
Recently computer vision has seen advancements mainly thanks to Transformer-based models. However many non-Transformer methods are still doing well being a direct competition of Transformer-based models. This review tries to present a comprehensive taxonomy of such methods and organize these methods into categories like convolution-based models, MLP-based models, state-space-based and more. These methods are looked at in terms of how efficient they are, how well they scale, how easy they are to understand and how robust they are. A total of 40 papers were chosen for this study. The goal is to give a view of non-Transformer methods and find out what challenges and opportunities exist for future computer vision research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。