arXiv:2605.14927cs.LG2026-05

研究输入特征的聚类结构如何影响浅层网络学习效率。

Learning with Shallow Neural Networks on Cluster-Structured Features

论文配图:Learning with Shallow Neural Networks on Cluster-Structured Features
图 1 · 摘自论文原文
  • 假设目标依赖少量隐含布尔变量,输入分组且与变量相关。
  • 梯度下降样本复杂度仅随隐变量数增长,高信噪比下与输入维度无关。
  • 适用于图像、文本等具空间相关性的数据学习分析。

深度学习在高维数据中的成功常归因于真实数据中存在低维结构。传统理论多假设结构存在于目标函数并投影输入至低维子空间,但图像、文本或基因序列等数据在输入空间本身具有强空间相关性。本文提出一个可处理的模型,研究此类相关性对浅层神经网络梯度下降学习样本复杂度的影响。具体考虑目标依赖少量隐含布尔变量,且输入特征分组并关联这些变量。在可识别性假设下,我们证明一种逐层梯度下降变体的样本复杂度随隐变量数增长,当信噪比足够高时,与输入维度无关(仅对数项依赖)。我们在合成数据和真实数据上验证了理论结果。

原文摘要 · Abstract (English)

The success of deep learning in high-dimensional settings is often attributed to the presence of low-dimensional structure in real-world data. While standard theoretical models typically assume that this structure lies in the target function, projecting unstructured inputs onto a low-dimensional subspace, data such as images, text or genomic sequences exhibit strong spatial correlations within the input space itself. In this paper, we propose a tractable model to study how these correlations affect the sample complexity of learning with gradient descent on shallow neural networks. Specifically, we consider targets that depend on a small number of latent Boolean variables, and input features grouped into clusters and correlated with the latent variables. Under an identifiability assumption, we show that for a layerwise gradient-descent variant, the sample complexity scales with the number of hidden variables and, when the signal-to-noise ratio is sufficiently high, is independent of the input dimension, up to logarithmic terms. We empirically test our theoretical findings on both synthetic and real data.

神经网络学习复杂度聚类结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。