神经网络在理论上能捕捉变量间真实关系,而传统模型可能学成常数预测。
Non-identifiability distinguishes Neural Networks among Parametric Models
- 通过不可识别性分析,揭示神经网络与传统参数模型的本质差异
- 在存在关系时,神经网络总能学到非平凡的映射;传统模型可能退化为均值预测
- 适用于理解深度学习为何比统计模型更擅长建模复杂依赖关系
神经网络与传统统计模型的核心区别长期悬而未决。本文在回归任务的总体层面证明:对任意随机变量对 $(X,Y)$,若存在真实关系,神经网络始终能学习到非平凡的映射。而对合理光滑的参数模型,在局部和全局可识别性条件下,存在非平凡的 $(X,Y)$ 对,使得该模型仅能学习常数预测 $\b{E}[Y]$。结果表明,不可识别性是区分神经网络与光滑参数模型的关键特征。
原文摘要 · Abstract (English)
One of the enduring problems surrounding neural networks is to identify the factors that differentiate them from traditional statistical models. We prove a pair of results which distinguish feedforward neural networks among parametric models at the population level, for regression tasks. Firstly, we prove that for any pair of random variables $(X,Y)$, neural networks always learn a nontrivial relationship between $X$ and $Y$, if one exists. Secondly, we prove that for reasonable smooth parametric models, under local and global identifiability conditions, there exists a nontrivial $(X,Y)$ pair for which the parametric model learns the constant predictor $\mathbb{E}[Y]$. Together, our results suggest that a lack of identifiability distinguishes neural networks among the class of smooth parametric models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。