arXiv:2510.07202cs.LG2025-10被引 1

研究深度窄网络逼近性能,揭示宽度临界点附近的失效机制

An in-depth look at approximation via deep and narrow neural networks

  • 在宽度w=n和w=n+1时逼近特定反例函数
  • 深度增加时逼近精度先升后降,出现性能瓶颈
  • 发现神经元死亡是导致性能下降的关键原因

2017年,Hanin和Sellke证明:任意深度、实值、前馈、ReLU激活的神经网络,当宽度w>n时,其函数类在紧集上一致收敛意义下稠密。为证明必要性,构造了一个具体反例函数f:R^n→R。本文深入研究该函数在宽度w=n与w=n+1时的逼近表现,分析不同深度下的逼近质量变化,并揭示(提示:神经元死亡)导致性能退化的机制。

原文摘要 · Abstract (English)

In 2017, Hanin and Sellke showed that the class of arbitrarily deep, real-valued, feed-forward and ReLU-activated networks of width w forms a dense subset of the space of continuous functions on R^n, with respect to the topology of uniform convergence on compact sets, if and only if w>n holds. To show the necessity, a concrete counterexample function f:R^n->R was used. In this note we actually approximate this very f by neural networks in the two cases w=n and w=n+1 around the aforementioned threshold. We study how the approximation quality behaves if we vary the depth and what effect (spoiler alert: dying neurons) cause that behavior.

神经网络逼近理论深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。