arXiv:2510.03511cs.CVcs.AI2025-10被引 8

让Transformer同时具备平移与多面体对称性,且不增加计算开销。

Platonic Transformers: A Solid Choice For Equivariance

  • 基于正多面体对称群定义注意力参考系,实现几何对称性约束
  • 在多个基准上性能媲美传统模型,且计算成本不变
  • 适合需要几何先验的视觉、分子预测任务

尽管广泛使用,但Transformer缺乏科学和计算机视觉中常见的几何对称性归纳偏置。现有等变方法通常通过复杂且计算量大的设计牺牲了Transformer的高效性和灵活性。我们提出Platonic Transformer,通过将注意力相对于正多面体对称群的参考系定义,建立了一种有原则的权值共享机制。该方法实现了连续平移与正多面体对称性的联合等变性,同时保持标准Transformer的完全架构和计算开销。此外,我们证明这种注意力在形式上等价于动态群卷积,揭示模型学习自适应几何滤波器,并可扩展为线性时间的卷积变体。在计算机视觉(CIFAR-10)、3D点云(ScanObjectNN)和分子性质预测(QM9, OMol25)等多个基准上,Platonic Transformer通过利用这些几何约束,在无额外开销下取得具有竞争力的性能。

原文摘要 · Abstract (English)

While widespread, Transformers lack inductive biases for geometric symmetries common in science and computer vision. Existing equivariant methods often sacrifice the efficiency and flexibility that make Transformers so effective through complex, computationally intensive designs. We introduce the Platonic Transformer to resolve this trade-off. By defining attention relative to reference frames from the Platonic solid symmetry groups, our method induces a principled weight-sharing scheme. This enables combined equivariance to continuous translations and Platonic symmetries, while preserving the exact architecture and computational cost of a standard Transformer. Furthermore, we show that this attention is formally equivalent to a dynamic group convolution, which reveals that the model learns adaptive geometric filters and enables a highly scalable, linear-time convolutional variant. Across diverse benchmarks in computer vision (CIFAR-10), 3D point clouds (ScanObjectNN), and molecular property prediction (QM9, OMol25), the Platonic Transformer achieves competitive performance by leveraging these geometric constraints at no additional cost.

Transformer等变性几何先验点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。