arXiv:2605.28704cs.LG2026-05

研究浮点神经网络在不精确计算下的表达能力,揭示哪些激活函数仍能逼近任意函数。

Expressive Power of Floating-Point Neural Networks with Arbitrary Reduction Orders and Inexact Activation Implementations

  • 构建通用可区分性框架,分析浮点计算中输入对的分辨能力
  • 证明在合理误差范围内,多种常见激活函数可实现全函数表示
  • 适用于真实浮点执行环境,突破传统理论限制,适合关注模型可靠性研究者

现有神经网络表达力理论多基于精确实数运算,但实际神经网络运行于有限精度浮点数下,且存在依赖实现的执行语义。近期工作虽开始研究浮点神经网络表达力,但仅限于受限激活函数及理想化假设(如固定左到右归约顺序、正确舍入激活)。本文研究在更一般浮点执行语义下的表达力,包括任意归约顺序与有界ulp误差的不精确激活实现。我们引入通用可区分性框架,证明第一层能区分所有不同输入是实现全域表示的必要条件。该刻画揭示了广义激活实现无法通用表示的类别,扩展了如正确舍入余弦激活等孤立反例。进一步证明,在激活实现满足弱条件时,适当可区分性亦为充分。基于此框架,我们在显著更现实的浮点模型下,建立了$ anh$、$ ext{Sigmoid}$、$ ext{ReLU}$、$ ext{ELU}$、$ ext{SeLU}$、$ ext{GeLU}$、$ ext{Swish}$、$ ext{Mish}$及$ ext{sin}$等多种实用激活函数的全函数表示性结果。

原文摘要 · Abstract (English)

Most existing expressivity theories for neural networks assume exact real arithmetic, whereas practical neural networks are executed under finite-precision floating-point arithmetic with implementation-dependent execution semantics. Recent works have begun studying the expressive power of floating-point neural networks, but existing results are limited to highly restricted activation functions and idealized assumptions such as fixed left-to-right reduction orders and correctly rounded activation implementations. In this work, we study the expressive power of floating-point neural networks under generalized floating-point execution semantics, including arbitrary reduction orders and inexact activation implementations with bounded ulp errors. We investigate when floating-point neural networks can represent arbitrary functions between floating-point domains exactly. To this end, we introduce a general distinguishability framework and show that the ability to distinguish every pair of distinct inputs in the first layer is necessary for universal representability. This characterization yields broad classes of activation implementations that are not universal representators, extending previous isolated counterexamples such as the correctly rounded cosine activation. We further prove that a suitable form of distinguishability is also sufficient for universal representability under mild conditions on the activation implementation. Using this framework, we establish universal representability results for a broad class of practical activation functions, including implementations of $\mathrm{Sigmoid}$, $\tanh$, $\mathrm{ReLU}$, $\mathrm{ELU}$, $\mathrm{SeLU}$, $\mathrm{GeLU}$, $\mathrm{Swish}$, $\mathrm{Mish}$, and $\sin$, under significantly more realistic floating-point execution models than previously known.

神经网络浮点计算表达力激活函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。