提出多尺度单指标模型,解析深度网络如何分层学习多尺度特征。
The Multiscale Single-Index Model: A Stylized Model for Hierarchical Feature Learning

- 构建分层特征提取的简化模型,每层在特定尺度上提取共享特征。
- 证明在线SGD仅需约d^{K-1}样本即可近乎完全恢复目标特征。
- 揭示深度优势与高阶统计结构的关系,适用于分析深层网络性能。
本文研究多尺度单指标模型(MSIM),该模型由 \\cite{oymak2021learning} 首次提出,用于刻画具有尺度分离特性的层级学习。每一层在某一物理尺度上提取共享的单指标特征并传递至下一层,形成可分析的深度架构学习框架。在链接函数非退化、植入选项非局域化假设下,目标函数的首阶维纳混沌呈现为受扰的尖刺张量,扰动大小为 d^{-1/2},使其成为张量PCA模型的自然非线性类比。虽然此微扰图像足以支持基于张量展开的高效谱恢复,但不足以分析基于反向传播的梯度方法。本文通过埃杰沃斯展开对维纳混沌进行精细分析:在一阶混沌中,得到尺度为 d^{-q/2} 的有限秩层次结构;在高阶混沌中,在自然非抵消条件下,平衡展开展现出大小为 d^{-ρ/2}、重数为 d^ρ 的阶梯式奇异值平台。利用这一高阶结构,并在额外慢赫米特能量尾部条件下,我们首先建立浅层网络近似下界,量化了深度带来的收益。更重要的是,证明在线SGD在相关性目标上,所有层以相同时间尺度演化时,可在 n = \widetilde{O}(d^{K-1}) 样本下实现 1 - o_d(1) 的恢复精度,达到与线性情形相同的样本复杂度。
原文摘要 · Abstract (English)
We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}. Each layer extracts a shared single-index feature at one physical scale and passes it to the next, thus defining a tractable setting in which to study how deep architectures learn multiscale representations. Under non-degeneracy and delocalization assumptions on the link function and planted features respectively, for fixed depth $K$ and local scale $d$, the first Wiener chaos of the target behaves as a perturbed spiked tensor, where the perturbation of order $d^{-1/2}$ comes from the non-linearity -- revealing the MSIM as a natural non-linear analogue of the Tensor PCA model \cite{montanari2014statistical}. While this perturbative picture is sufficient to enable efficient spectral recovery based on Tensor unfolding (as already observed in \cite{oymak2021learning}), it is not precise enough for the analysis of backpropagation gradient-based methods. In this work, we address this limitation by performing a fine-grained analysis of the Wiener chaos using Edgeworth expansions. In the first chaos, this gives a finite-rank hierarchy at scales $d^{-q/2}$. In higher chaoses, balanced flattenings exhibit staircase singular-value plateaus of size $d^{-ρ/2}$ and multiplicity $d^ρ$ under a natural higher-chaos non-cancellation condition. Using this higher-chaos structure, and under an additional slow Hermite-energy tail condition, we first establish shallow-network approximation lower bounds, quantifying the benefit of depth in this model. Next, and most importantly, we prove that online SGD on the correlation objective, where all layers evolve in the same timescale, achieves $1 - o_d(1)$ recovery with $n = \widetilde{O}( d^{K-1})$ samples, recovering the same sample complexity as in the linear counterpart.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。