arXiv:2609.06430cs.LGstat.ML2026-09

解释了二层量化网络中梯度估计的泛化能力,为训练稳定性提供理论支持。

Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks

  • 从统计学习角度分析阶梯式梯度估计的稳定性机制。
  • 证明了模型在样本上的平均稳定界为 $O(n^{-1/2})$,并给出风险上界。
  • 适用于关注量化神经网络理论性质的研究者与工程师。

本文从统计学习理论视角研究两层二值激活网络在铰链损失下的身份直通估计器(STE)。核心问题是:算法稳定性能否解释不连续STE训练规则产生的泛化性能?在输出饱和情形下,零初始化的样本级STE递推等价于凸潜空间损失 $(-yu^ op x)_+$ 的随机次梯度下降。该表示使稳定性分析成为可能。我们推导出两个耦合更新间的精确距离恒等式,证明公共样本映射近似非扩张性,当两个潜空间边界跨过零时存在二次偏差。由此获得显式的 $oldsymbol{ ext{ℓ}_2}$ 平均模型稳定性与泛化界,并将稳定性从潜向量等距传递至完整第一层矩阵。结合优化界,得到显式的过度诱导风险保证,在 $T=n^2$ 时达到 $O(n^{-1/2})$ 收敛率。在边界可分条件下,进一步给出随机单遍迭代的最优阶 $O(R^2/(oldsymbol{ ext{γ}}^2n))$ 期望超额误分类误差及相应多数投票界。

原文摘要 · Abstract (English)

We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss $(-yu^\top x)_+$. This representation makes a stability analysis possible. We derive an exact distance identity for two coupled updates and prove approximate non-expansiveness of the common-example map, with a quadratic defect only when the two latent margins straddle zero. We then obtain explicit $\ell_2$ on-average model-stability and generalization bounds, transferring stability isometrically from the latent vector to the full first-layer matrix. Combining stability with a standard optimization bound yields an explicit excess induced-risk guarantee and the rate $O(n^{-1/2})$ when $T=n^2$. Under margin separability, a complementary argument gives the optimal-order $O(R^2/(\gamma^2n))$ expected excess misclassification error for a randomized one-pass STE iterate and a corresponding majority-vote bound.

量化网络泛化理论梯度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。