通过布尔函数研究神经网络的归纳偏置与泛化关系
Characterising the Inductive Biases of Neural Networks on Boolean Data
- 用深度2全连接网络对应析取范式,解析建模归纳偏置
- 训练中特征可解释,且动态规律可预测
- 适合关注模型泛化机制的研究者
深度神经网络在参数量远超数据量的情况下仍能良好泛化。现有研究仅部分解释此现象(如基于NTK的任务-模型对齐理论忽略了特征学习)。本文通过深度2离散全连接网络与析取范式(DNF)公式的双射关系,在布尔函数上开展端到端、可解析的案例研究,揭示了网络归纳先验、训练动态(包括特征学习)与最终泛化之间的联系。在蒙特卡洛学习算法下,模型展现出可预测的训练过程和可解释特征的涌现。该框架可细致追踪归纳偏置与特征形成如何驱动泛化。
原文摘要 · Abstract (English)
Deep neural networks are renowned for their ability to generalise well across diverse tasks, even when heavily overparameterized. Existing works offer only partial explanations (for example, the NTK-based task-model alignment explanation neglects feature learning). Here, we provide an end-to-end, analytically tractable case study that links a network's inductive prior, its training dynamics including feature learning, and its eventual generalisation. Specifically, we exploit the one-to-one correspondence between depth-2 discrete fully connected networks and disjunctive normal form (DNF) formulas by training on Boolean functions. Under a Monte Carlo learning algorithm, our model exhibits predictable training dynamics and the emergence of interpretable features. This framework allows us to trace, in detail, how inductive bias and feature formation drive generalisation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。