深度学习模型因目标函数不确定和对称性,存在海量等效最优解。
Permutative redundancy and uncertainty of the objective in deep learning
- 利用对称性与目标不确定性分析模型优化难题
- 网络规模增大时全局最优解形成复杂峡谷与山脊结构
- 提出剪枝、重排、正交激活等方法消除冗余解
本文探讨了目标函数不确定性和传统深度学习架构的置换对称性带来的影响。研究表明,传统架构被海量等效全局与局部最优解污染。目标函数的不确定性导致局部最优无法达到,且随着网络规模增大,全局优化景观可能演变为复杂的峡谷与山脊交织结构。文中讨论了若干缓解或消除虚假最优解的方法,包括强制预剪枝、重新排序、正交多项式激活以及模块化生物启发式架构。
原文摘要 · Abstract (English)
Implications of uncertain objective functions and permutative symmetry of traditional deep learning architectures are discussed. It is shown that traditional architectures are polluted by an astronomical number of equivalent global and local optima. Uncertainty of the objective makes local optima unattainable, and, as the size of the network grows, the global optimization landscape likely becomes a tangled web of valleys and ridges. Some remedies which reduce or eliminate ghost optima are discussed including forced pre-pruning, re-ordering, ortho-polynomial activations, and modular bio-inspired architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。