arXiv:2602.10195cs.LGcs.AI2026-02被引 1

用几何代数构建新序列模型,性能超Transformer且更可解释。

Versor: A Geometric Sequence Architecture

  • 用旋量在几何流形上演化状态,天然保持三维对称性。
  • 参数量少200倍,零样本扩展能力远超ViT(MCC 0.993 vs 0.070)。
  • 适合需要高效、可解释建模的物理系统与复杂关系任务。

提出新型序列架构Versor,采用共形几何代数(CGA)替代传统线性操作,在混沌多体动力学、拓扑推理及标准多模态基准(CIFAR-10、WikiText-103)上持续超越Transformer、图网络及几何基线(GATr、EGNN)。关键成果包括:参数量减少200倍;注意力可分解为距离与方向成分,具备可解释性;零样本尺度泛化能力优异(MCC 0.993 vs ViT的0.070);引入递归旋量累加器(RRA),实现动力系统下O(L)线性时间复杂度;几何乘积注意力(GPA)支持O(L²)全局关系建模,支持按任务需求剪枝或混合架构。分布外测试中,Versor预测稳定,而Transformer崩溃。定制克莱夫德核通过位掩码收缩与专用矩阵同构内核,实现超100倍加速,单步延迟降至1.05毫秒,优于高度优化的Transformer基线。

原文摘要 · Abstract (English)

A novel sequence architecture is introduced, Versor, which uses Conformal Geometric Algebra (CGA) in place of traditional linear operations to achieve structural generalization and significant performance improvements on a variety of tasks, while offering improved interpretability and efficiency. By embedding states in the $Cl_{4,1}$ manifold and evolving them via geometric transformations (rotors), Versor natively represents $SE(3)$-equivariant relationships without requiring explicit structural encoding. Versor is validated on chaotic N-body dynamics, topological reasoning, and standard multimodal benchmarks (CIFAR-10, WikiText-103), consistently outperforming Transformers, Graph Networks, and geometric baselines (GATr, EGNN). Key results include: orders-of-magnitude fewer parameters ($200\times$ vs. Transformers); interpretable attention decomposing into proximity and orientational components; zero-shot scale generalization (0.993 vs. 0.070 MCC for ViT); and featuring a Recursive Rotor Accumulator (RRA) for $O(L)$ linear temporal complexity in dynamical systems, and a Geometric Product Attention (GPA) mechanism for $O(L^{2})$ global relational modeling, allowing for task-specific architectural pruning or hybridization depending on the required scale. In out-of-distribution tests, Versor maintains stable predictions while Transformers fail catastrophically. Custom Clifford kernels achieve a cumulative over $100\times$ speedup via bit-masked contraction and specialized Matrix Isomorphism kernels, reducing per-step latency to 1.05 ms and outperforming highly-optimized Transformer baselines.

几何代数序列建模可解释性高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。