arXiv:2412.09779stat.MLcs.LG2024-12被引 2

提出新维度定义,让深度学习在低维数据上收敛更快更准。

A Statistical Analysis for Supervised Deep Learning with Exponential Families for Intrinsically Low-dimensional Data

  • 用熵定义数据内在维度,替代传统几何维度
  • 理论证明误差随样本数呈多项式衰减,不指数增长
  • 适合研究低维数据中指数族分布的深度学习建模

近期研究表明,深度监督学习的期望测试误差收敛速率取决于数据的内在维度,而非输入空间的维数 $d$。现有文献将内在维度定义为支撑集的闵可夫斯基维数或流形维数,常导致次优率和不切实际的假设。本文考虑响应变量在解释变量给定下服从指数族分布且均值函数为 $β$-霍尔德光滑的情形。引入解释变量分布 $λ$ 的 $2β$-熵维 $ar{d}_{2β}(λ)$ 概念,证明在 $n$ 个独立同分布样本下,测试误差为 $ ilde{oldsymbol{ ext{O}}}ig(n^{- rac{2β}{2β+ ar{d}_{2β}(λ)}}ig)$,优于已有最优率。进一步,在解释变量密度有界条件下,收敛率可表为 $ ilde{oldsymbol{ ext{O}}}ig( d^{ rac{2 loor{β}(β+ d)}{2β+ d}} n^{- rac{2β}{2β+ d}} ig)$,表明 $d$ 的依赖关系至多为多项式,非指数级。当解释变量密度有下界时,该样本依赖率对学习指数族依赖结构几乎最优。

原文摘要 · Abstract (English)

Recent advances have revealed that the rate of convergence of the expected test error in deep supervised learning decays as a function of the intrinsic dimension and not the dimension $d$ of the input space. Existing literature defines this intrinsic dimension as the Minkowski dimension or the manifold dimension of the support of the underlying probability measures, which often results in sub-optimal rates and unrealistic assumptions. In this paper, we consider supervised deep learning when the response given the explanatory variable is distributed according to an exponential family with a $β$-Hölder smooth mean function. We consider an entropic notion of the intrinsic data-dimension and demonstrate that with $n$ independent and identically distributed samples, the test error scales as $\tilde{\mathcal{O}}\left(n^{-\frac{2β}{2β+ \bar{d}_{2β}(λ)}}\right)$, where $\bar{d}_{2β}(λ)$ is the $2β$-entropic dimension of $λ$, the distribution of the explanatory variables. This improves on the best-known rates. Furthermore, under the assumption of an upper-bounded density of the explanatory variables, we characterize the rate of convergence as $\tilde{\mathcal{O}}\left( d^{\frac{2\lfloorβ\rfloor(β+ d)}{2β+ d}}n^{-\frac{2β}{2β+ d}}\right)$, establishing that the dependence on $d$ is not exponential but at most polynomial. We also demonstrate that when the explanatory variable has a lower bounded density, this rate in terms of the number of data samples, is nearly optimal for learning the dependence structure for exponential families.

深度学习统计分析指数族低维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。