arXiv:2608.06687math.NAcs.LG2026-08

用确定采样点优化神经网络逼近椭圆谱方程,理论证明误差随参数量下降。

Optimal Neural Network Approximation via Empirical Least Squares with Deterministic Samples

  • 基于确定性采样点的离散最小二乘法逼近椭圆谱方程。
  • 当参数数n增长时,逼近误差以n^{-r/d}速率收敛,依赖于函数光滑性。
  • 理论适用于球面上的ReLU^k网络,适合研究神经网络逼近的数学基础者。

我们为椭圆谱方程 $\mathfrak L_βu=f$ 构建了基于线性化 ReLU$^k$ 神经网络在球面 $\mathbb S^d$ 上的离散残差最小二乘逼近的严格理论,其中 $\mathfrak L_β$ 是阶为 $β$ 的正椭圆谱乘子。给定参数集 $Θ_n=\{θ_{j}^*\}_{j=1}^n\subset\mathbb S^d$,在列点 $\{η_i^*\}_{i=1}^m$ 处通过最小化残差平方和来逼近解 $u$:$u_{n,m}\in\arg\min_{v_n\in L_n^k(Θ_n)}\frac1m\sum_{i=1}^m\left(f(η_i^*)-\mathfrak L_βv_n(η_i^*)\right)^2$。当 $k>\frac{d-1}{2}+β$,且网络参数为反极准均匀、列点为任意准均匀且 $m\gtrsim n$ 时,证明了误差范数 $\|u-u_{n,m}\|_{\mathcal H^β(\mathbb S^d)}\eqsim\|f-\mathfrak L_βu_{n,m}\|_{\mathcal L^2(\mathbb S^d)}\lesssim n^{-\frac{r}{d}}$,其中若 $\frac{d}{p}<r\leq \frac{d}{2}, p>2$,则上界为 $\|f\|_{\mathcal W^{r,p}(\mathbb S^d)}$;若 $r>\frac{d}{2}$,则为 $\|f\|_{\mathcal H^r(\mathbb S^d)}$。还建立了对独立同分布均匀列点的高概率残差估计,误差中含对数因子与任意小的光滑性损失。关键分析工具是线性化 ReLU$^k$ 网络空间的 Bernstein 不等式:若 $\underline h$ 表示参数的反极分离距离,则 $\|v_n\|_{\mathcal H^r(\mathbb S^d)}\lesssim\underline h^{-(r-s)}\|v_n\|_{\mathcal H^s(\mathbb S^d)}$,其中 $0\leq s<r<k+\tfrac12$。

原文摘要 · Abstract (English)

We develop a rigorous theory of discrete residual least-squares approximation for elliptic spectral equations $\mathfrak L_βu=f$ using linearized ReLU$^k$ neural networks on the sphere, where $\mathfrak L_β$ is a positive elliptic spectral multiplier of order $β$. Given a parameter set $Θ_n=\{θ_{j}^*\}_{j=1}^n\subset\mathbb S^d$, we approximate $u$ in the linearized network space $L_n^k(Θ_n)$ by the discrete residual on the collocation points $\{η_i^*\}_{i=1}^m$ \begin{equation*} u_{n,m}\in\arg\min_{v_n\in L_n^k(Θ_n)}\frac1m\sum_{i=1}^m\left(f(η_i^*)-\mathfrak L_βv_n(η_i^*)\right)^2. \end{equation*} With $k>\frac{d-1}{2}+β$, for antipodally quasi-uniform network parameter sets and any quasi-uniform collocation points with $m\gtrsim n$, we prove that \begin{equation*} \|u-u_{n,m}\|_{\mathcal H^β(\mathbb S^d)}\eqsim\|f-\mathfrak L_βu_{n,m}\|_{\mathcal L^2(\mathbb S^d)}\lesssim n^{-\frac{r}{d}} \begin{cases} \|f\|_{\mathcal W^{r,p}(\mathbb S^d)},&\frac{d}{p}<r\leq \frac{d}{2},~p>2,\\ \|f\|_{\mathcal H^r(\mathbb S^d)},&r>\frac{d}{2}. \end{cases} \end{equation*} We also establish a high-probability residual estimate, up to a logarithmic factor and an arbitrarily small smoothness loss, for i.i.d.\ uniformly distributed collocation points. The key analytical ingredient is a Bernstein inequality for linearized ReLU$^k$ network spaces. If $\underline h$ denotes the antipodal separation distance of the network parameters, then \begin{equation*} \|v_n\|_{\mathcal H^r(\mathbb S^d)}\lesssim\underline h^{-(r-s)}\|v_n\|_{\mathcal H^s(\mathbb S^d)},\qquad 0\leq s<r<k+\tfrac12. \end{equation*}

神经网络逼近理论椭圆方程球面

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。