arXiv:2411.15095stat.MLcs.CV2024-11被引 5

深度神经网络在结构化密度估计中实现与维度无关的收敛速度。

Dimension-independent rates for structured neural density estimation

  • 用简单L²损失训练神经网络,学习马尔可夫图最大团大小为r的密度分布。
  • 在真实图像、音频等数据中,团大小通常为常数(r=O(1)),收敛率达n^{-1/(4+r)}。
  • 相比传统非参数方法,有效维度由最大团大小决定,适合高维数据建模。

我们证明深度神经网络在学习图像、音频、视频和文本等应用中的结构化密度时,能达到与维度无关的收敛速率。具体而言,当底层密度服从最大团大小不超过r的马尔可夫随机场时,采用简单L²最小化损失的神经网络在非参数密度估计中可达到n^{-1/(4+r)}的收敛率。我们提供证据表明,在上述应用中,该团大小通常为常数(即r=O(1))。进一步,我们建立了L¹最优收敛率为n^{-1/(2+r)},相较于标准非参数速率n^{-1/(2+d)},揭示此类问题的有效维度是马尔可夫随机场的最大团大小。这些速率独立于数据的环境维度,适用于真实场景下的图像、声音、视频和文本模型。结果为深度学习克服维度灾难提供了新解释,展示了在这些场景下维度无关的收敛性。

原文摘要 · Abstract (English)

We show that deep neural networks achieve dimension-independent rates of convergence for learning structured densities such as those arising in image, audio, video, and text applications. More precisely, we demonstrate that neural networks with a simple $L^2$-minimizing loss achieve a rate of $n^{-1/(4+r)}$ in nonparametric density estimation when the underlying density is Markov to a graph whose maximum clique size is at most $r$, and we provide evidence that in the aforementioned applications, this size is typically constant, i.e., $r=O(1)$. We then establish that the optimal rate in $L^1$ is $n^{-1/(2+r)}$ which, compared to the standard nonparametric rate of $n^{-1/(2+d)}$, reveals that the effective dimension of such problems is the size of the largest clique in the Markov random field. These rates are independent of the data's ambient dimension, making them applicable to realistic models of image, sound, video, and text data. Our results provide a novel justification for deep learning's ability to circumvent the curse of dimensionality, demonstrating dimension-independent convergence rates in these contexts.

密度估计深度学习维度灾难马尔可夫随机场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。