arXiv:2608.31157cs.LGstat.ML2026-08

研究低维隐变量如何高效逼近大模型参数,揭示压缩与精度的数学极限。

Sharp Approximation Rates for Neural Networks with Affine Latent Parameterizations

  • 用仿射生成器将低维隐变量映射为网络参数,统一多种高效方法
  • 证明在 $P\min\{M,P\}$ 个参数下,逼近误差为 $ (P\min\{M,P\})^{-\alpha/d} $
  • 即使隐空间固定维度,增大网络规模也能逼近任意精度,适合模型压缩研究

许多参数高效方法通过低维隐表示生成大型神经网络的参数。给定一个具有 $P_\Phi$ 个参数槽的架构 $\Phi$,我们记 $\boldsymbol{\theta}_f=\mathcal{G}(\boldsymbol{\xi}_f)$,其中 $\mathcal{G}:\mathbb{R}^M\to\mathbb{R}^{P_\Phi}$ 是参数生成器,$\boldsymbol{\xi}_f\in\mathbb{R}^M$ 是目标函数 $f$ 的隐表示。架构 $\Phi$ 和生成器 $\mathcal{G}$ 在整个目标类中共享,每个目标 $f$ 由其独立的隐向量 $\boldsymbol{\xi}_f$ 表示,使得 $\Phi_{\mathcal{G}(\boldsymbol{\xi}_f)}$ 近似 $f$。该框架涵盖超网络、低维参数化、参数高效微调和模型压缩。理解隐维数 $M$ 与网络预算 $P$ 之间的权衡对刻画这些方法的表达效率至关重要。本文针对仿射生成器和全连接 ReLU 架构,研究此权衡。更精确地,在满足 $P_\Phi\leq P$ 的架构 $\Phi$ 和仿射生成器 $\mathcal{G}:\mathbb{R}^M\to \mathbb{R}^{P_\Phi}$ 上联合优化,证明单位球上 $\alpha$-H"older 函数($0<\alpha\leq1$)的最优最坏情况一致逼近误差具有尖锐阶 $ (P\min\{M,P\})^{-\alpha/d} $。特别地,结果表明即使隐空间维度固定,随着网络预算增加,逼近误差仍可趋于零。

原文摘要 · Abstract (English)

Many parameter-efficient methods generate the parameters of a large neural network from a low-dimensional latent representation. Given an architecture $Φ$ with $P_Φ$ parameter slots, we write $\boldsymbolθ_f=\mathcal{G}(\boldsymbolξ_f)$, where $\mathcal{G}\colon\mathbb{R}^M\to\mathbb{R}^{P_Φ}$ is a parameter generator and $\boldsymbolξ_f\in\mathbb{R}^M$ is a latent representation of the target function $f$. The architecture $Φ$ and the generator $\mathcal{G}$ are shared across the entire target class, while each target $f$ is represented by its own latent vector $\boldsymbolξ_f$, with $Φ_{\mathcal{G}(\boldsymbolξ_f)}$ approximating $f$. This framework encompasses hypernetworks, low-dimensional parameterizations, parameter-efficient adaptation, and model compression. Understanding the tradeoff between the latent dimension $M$ and the network budget $P$ is therefore fundamental to characterizing the expressive efficiency of these methods. We study this tradeoff for affine generators and fully connected ReLU architectures. More precisely, optimizing jointly over architectures $Φ$ satisfying $P_Φ\leq P$ and affine generators $\mathcal{G}:\mathbb{R}^M\to \mathbb{R}^{P_Φ}$, we prove that the optimal worst-case uniform approximation error over the unit ball of $α$-Hölder functions on $[0,1]^d$, where $0<α\leq1$, has the sharp order $ \bigl(P\min\{M,P\}\bigr)^{-α/d}. $ In particular, our result shows that even a fixed-dimensional latent space suffices to achieve vanishing approximation error as the network budget increases.

神经网络参数效率逼近理论模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。