提出RAF模型,解析神经网络如何同时学规律与记例外。
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
- 用结构化规则+随机例外构造训练数据,建模双能力
- 过参数化支持记忆,正则化调控规则与记忆的平衡
- 适合研究模型泛化与记忆机制的理论工作者
现代神经网络兼具学习底层规律和记忆特定事实或例外的能力,但其理论理解仍不充分。本文提出规则与事实(RAF)模型,一个最小可解框架,通过连接统计学习物理中的教师-学生模型与格登式容量分析,精确刻画该现象。模型中,$1 - \varepsilon$ 的训练标签由结构化教师规则生成,$\varepsilon$ 为随机标签的非结构化事实。我们刻画了学习器在新数据上泛化并记忆非结构化样本的条件。结果表明:足够的过参数化支持记忆,而正则化、核函数或非线性选择决定容量在规则学习与记忆之间的分配。RAF模型为理解现代神经网络如何在存储稀有或不可压缩信息的同时推断结构提供了理论基础。
原文摘要 · Abstract (English)
A key capability of modern neural networks is their capacity to simultaneously learn underlying rules and memorize specific facts or exceptions. Yet, theoretical understanding of this dual capability remains limited. We introduce the Rules-and-Facts (RAF) model, a minimal solvable setting that enables precise characterization of this phenomenon by bridging two classical lines of work in the statistical physics of learning: the teacher-student framework for generalization and Gardner-style capacity analysis for memorization. In the RAF model, a fraction $1 - \varepsilon$ of training labels is generated by a structured teacher rule, while a fraction $\varepsilon$ consists of unstructured facts with random labels. We characterize when the learner can simultaneously recover the underlying rule - allowing generalization to new data - and memorize the unstructured examples. Our results quantify how overparameterization enables the simultaneous realization of these two objectives: sufficient excess capacity supports memorization, while regularization and the choice of kernel or nonlinearity control the allocation of capacity between rule learning and memorization. The RAF model provides a theoretical foundation for understanding how modern neural networks can infer structure while storing rare or non-compressible information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。