深度学习的泛化现象并非神秘,可用经典理论解释。
Deep Learning is Not So Mysterious or Different
- 用软归纳偏置统一解释过拟合、双下降等现象
- 在经典框架下证明深度学习可被严谨分析
- 适合想理解深度学习本质的研究者
深度神经网络常被认为与传统模型不同,表现出反直觉的泛化行为,如良性过拟合、双下降和过参数化的成功。我们主张这些现象并非神经网络独有,也不特别神秘。通过长期存在的泛化框架(如PAC-Bayes和可数假设界),这些行为既可直观理解,又能严格刻画。核心在于‘软归纳偏置’:不强行限制假设空间,而是对与数据一致的简单解施加柔性偏好。该原则适用于多种模型类,表明深度学习并不像表面看起来那般神秘或特殊。但我们也指出,深度学习在表征学习、模式连通性及通用性方面仍具独特性。
原文摘要 · Abstract (English)
Deep neural networks are often seen as different from other model classes by defying conventional notions of generalization. Popular examples of anomalous generalization behaviour include benign overfitting, double descent, and the success of overparametrization. We argue that these phenomena are not distinct to neural networks, or particularly mysterious. Moreover, this generalization behaviour can be intuitively understood, and rigorously characterized, using long-standing generalization frameworks such as PAC-Bayes and countable hypothesis bounds. We present soft inductive biases as a key unifying principle in explaining these phenomena: rather than restricting the hypothesis space to avoid overfitting, embrace a flexible hypothesis space, with a soft preference for simpler solutions that are consistent with the data. This principle can be encoded in many model classes, and thus deep learning is not as mysterious or different from other model classes as it might seem. However, we also highlight how deep learning is relatively distinct in other ways, such as its ability for representation learning, phenomena such as mode connectivity, and its relative universality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。