arXiv:2606.27864cs.CVcs.LG2026-06

提出统一框架,让视觉Transformer具备平面离散对称性不变性。

A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of $\mathrm{O}(2)$

论文配图:A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of $\mathrm{O}(2)$
图 1 · 摘自论文原文
  • 构建可适配任意O(2)离散子群的等变Transformer架构
  • 在PatternNet数据集上验证,等变性提升识别准确率
  • 支持六重旋转对称,适用于航空图像等对称场景

视觉变换器已成为视觉识别的主流架构,但标准模型未显式编码许多视觉任务中的平面对称性。本文提出一族对任意O(2)离散子群等变的视觉变换器,构建统一框架,推广了先前的翻转与D₄等变架构。该构造生成核心变换器组件的等变版本,并提供表达能力保证。特别地,当H ≤ G时,G-等变ViT类自然嵌入H-等变ViT类。在单头设置下,对应的等变自注意力层能实现所有由普通自注意力表示的G-等变自注意力映射。我们还基于六边形块构建了D₆等变模型,兼容六重旋转对称性。在PatternNet航空图像数据集上,于人为数据稀缺条件下评估了不同子群(D₄与D₆)下的模型表现。实验比较两种等变注意力机制,并分析非线性中同质空间配置选择对性能的影响。初步结果在参数量匹配情况下显示,等变性有助于提升识别准确率,推动进一步研究离散对称群如何塑造基于变换器的视觉识别模型。

原文摘要 · Abstract (English)

Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetries that arise in many vision domains. We introduce a family of vision transformers equivariant to arbitrary discrete subgroups of $\mathrm{O}(2)$, providing a unified framework that generalizes prior flipping- and $D_4$-equivariant transformer architectures. Our construction yields equivariant analogues of the core transformer components, together with expressivity guarantees for the resulting layers. In particular, we show that whenever $H \le G$, the class of $G$-equivariant ViTs embeds naturally into the class of $H$-equivariant ViTs. We also prove that, in the single-head setting, the corresponding equivariant self-attention layer realizes every $G$-equivariant self-attention map representable by ordinary self-attention. We further construct a $D_6$-equivariant model based on hexagonal patches, making the architecture compatible with six-fold rotational symmetries. We evaluate the resulting models on the PatternNet aerial image dataset in artificially data-scarce regimes across subgroups of $D_4$ and $D_6$. Our experiments compare two equivariant attention mechanisms and analyze how the choice of homogeneous-space configurations used in the nonlinearities affects performance. Preliminary results under matched parameter budgets indicate that equivariance can improve recognition accuracy, motivating further study of how discrete symmetry groups shape transformer-based visual recognition models.

视觉Transformer等变神经网络对称性建模图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。