提出一种新型视觉变换器,能自动处理图像旋转翻转对称性,提升分类准确率。
REViT: Roto-reflection Equivariant Convolutional Vision Transformer

- 用卷积注意力设计旋转翻转等变的视觉变换器
- 在图像分类任务上超越现有等变网络方法
- 适合需要方向敏感性的视觉识别场景
本文提出一种基于卷积注意力的离散旋转变换群等变视觉变换器。该模型保持特征图中的旋转、翻转和位置对称性,适用于输入方向影响输出的任务。尽管现有研究多聚焦于卷积神经网络实现等变性,但本工作探讨了在视觉变换器中实现等变性的挑战,并提出一种更简单的离散旋转变换群等变视觉变换器构建方式。实验结果表明,该方法在图像分类任务中优于现有的离散旋转变换群等变神经网络方案。
原文摘要 · Abstract (English)
In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs. In image classification and object detection, most of the studies on roto-reflection equivariant models have focused on using convolutional neural networks rather than vision transformers. In this paper, we examine the challenges involved in achieving equivariance in vision transformers, and we propose a simpler way to implement a discretized roto-reflection group equivariant vision transformer. The experimental results demonstrate that our approach outperforms the existing approaches for developing discrete roto-reflection group equivariant neural networks for image classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。