arXiv:2604.24878cs.LGcs.AI2026-04

用ReLU近似方法解析Transformer注意力机制,给出高效计算资源边界。

Transformer Approximations from ReLUs

  • 将ReLU近似技术系统性转化为软注意力机制的构造方法
  • 实现乘法、倒数、极值等运算的精准低资源近似
  • 为分析Transformer模型提供新理论工具,适合算法优化研究者

我们提供了一种系统性方法,将ReLU近似结果转化为软注意力机制的构造方案。该方法适用于多种常见近似目标,重要的是,它能给出针对特定目标的经济高效资源约束,超越泛化近似结论。我们在乘法、倒数计算和极值(min/max)等基本运算上展示了该方法的应用。这些成果为分析软注意力型Transformer模型提供了新的理论工具。

原文摘要 · Abstract (English)

We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation targets. Importantly, it yields target-specific, economic resource bounds beyond universal approximation statements. We showcase the recipe on multiplication, reciprocal computation, and min/max primitives. These results provide new analytical tools for analyzing softmax transformer models.

Transformer近似计算注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。