通过代数结构嵌入层级信息,提升模型效率与可解释性。
Graded Transformers
- 用可学习的分级变换构建注意力与表征层,编码层次结构。
- 理论证明具备通用逼近能力,样本复杂度更低且对扰动鲁棒。
- 适合需要层级推理的领域,如语言、生物序列与安全关键系统。
我们提出分级变压器(Graded Transformer)框架,一种通过向量空间上的分级变换嵌入代数归纳偏置的新序列模型。扩展自分级神经网络(GNNs),本文设计两种架构:线性分级变压器(LGT)和指数分级变压器(EGT)。它们在注意机制和表示层中应用参数化缩放算子,由固定或可学习的分级元组控制,在EGT中引入指数因子,以编码层级结构并提升对结构化数据的效率。我们建立了严格的理论保证,包括连续函数与Sobolev函数的通用逼近定理,通过有效VC维界降低样本复杂度,以及分级操作的Lipschitz连续性与抗扰动能力。分级损失确保优化过程中的梯度稳定性和与领域先验的一致性。通过将等级视为可微参数,该框架实现特征的自适应优先级分配,克服了早期模型中固定等级的局限。该框架为层级学习与神经符号推理提供了数学上严谨的方法,应用场景涵盖代数几何(模空间与黎曼ζ函数)、物理(多尺度系统)、自然语言处理(句法解析)、生物序列分析(变异预测)、机器人与自动驾驶系统(安全关键优先级)、汽车工业(可认证的ADAS AI)以及区块链与金融密码学(安全编码与结构化预测)。
原文摘要 · Abstract (English)
We introduce the Graded Transformer framework, a new class of sequence models that embeds algebraic inductive biases through grading transformations on vector spaces. Extending Graded Neural Networks (GNNs), we propose two architectures: the Linearly Graded Transformer (LGT) and the Exponentially Graded Transformer (EGT). These models apply parameterized scaling operators, governed by fixed or learnable grading tuples and in the case of EGT exponential factors, to encode hierarchical structure in attention and representation layers and to improve efficiency for structured data. We establish rigorous guarantees, including universal approximation theorems for continuous and Sobolev functions, reduced sample complexity via effective VC dimension bounds, Lipschitz continuity of graded operations, and robustness to perturbations. A graded loss ensures gradient stability and alignment with domain priors during optimization. By treating grades as differentiable parameters, the framework enables adaptive feature prioritization, overcoming limitations of fixed grades in earlier models. The Graded Transformer provides a mathematically principled approach to hierarchical learning and neuro-symbolic reasoning. Applications include algebraic geometry (moduli spaces and zeta functions), physics (multiscale systems), natural language processing (syntactic parsing), biological sequence analysis (variant prediction), robotics and autonomous systems (safety-critical prioritization), the automotive industry (certifiable AI for ADAS), and blockchain and financial cryptography (secure coding and structured prediction).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。