揭示深度神经网络可训练性与泛化能力的内在机制。
Conjugate Learning Theory: Uncovering the Mechanisms of Trainability and Generalization in Deep Neural Networks
- 基于对偶理论构建可训练性分析框架,解析SGD优化过程
- 证明小批量SGD能达经验风险全局最优,受结构矩阵特征值与梯度能量控制
- 首次给出模型无关的经验风险下界,揭示数据决定训练极限
本文提出一种基于有限样本的实用可训练性概念,构建了基于凸对偶理论的共轭学习理论框架。在此基础上,证明使用小批量随机梯度下降(SGD)训练深度神经网络(DNNs)时,通过联合调控结构矩阵的极值特征值与梯度能量,可实现经验风险的全局最优,并建立了相应收敛定理。进一步揭示了批大小、模型架构(包括深度、参数量、稀疏性、跳跃连接等)对非凸优化的影响。推导出模型无关的经验风险下界,理论上证明数据决定了训练性的根本极限。在泛化方面,基于广义条件熵提出了确定性和概率性泛化误差上界:前者明确泛化误差范围,后者在i.i.d.采样下刻画误差分布;两者均量化了三要素影响:(i) 模型不可逆性导致的信息损失,(ii) 可达最大损失值,(iii) 特征相对于标签的广义条件熵。该框架统一解释了正则化、不可逆变换与网络深度对泛化的作用。大量实验验证了所有理论预测,确认框架正确性与一致性。
原文摘要 · Abstract (English)
In this work, we propose a notion of practical learnability grounded in finite sample settings, and develop a conjugate learning theoretical framework based on convex conjugate duality to characterize this learnability property. Building on this foundation, we demonstrate that training deep neural networks (DNNs) with mini-batch stochastic gradient descent (SGD) achieves global optima of empirical risk by jointly controlling the extreme eigenvalues of a structure matrix and the gradient energy, and we establish a corresponding convergence theorem. We further elucidate the impact of batch size and model architecture (including depth, parameter count, sparsity, skip connections, and other characteristics) on non-convex optimization. Additionally, we derive a model-agnostic lower bound for the achievable empirical risk, theoretically demonstrating that data determines the fundamental limit of trainability. On the generalization front, we derive deterministic and probabilistic bounds on generalization error based on generalized conditional entropy measures. The former explicitly delineates the range of generalization error, while the latter characterizes the distribution of generalization error relative to the deterministic bounds under independent and identically distributed (i.i.d.) sampling conditions. Furthermore, these bounds explicitly quantify the influence of three key factors: (i) information loss induced by irreversibility in the model, (ii) the maximum attainable loss value, and (iii) the generalized conditional entropy of features with respect to labels. Moreover, they offer a unified theoretical lens for understanding the roles of regularization, irreversible transformations, and network depth in shaping the generalization behavior of deep neural networks. Extensive experiments validate all theoretical predictions, confirming the framework's correctness and consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。