用受限玻尔兹曼机研究教师-学生学习中的结构化数据建模。
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
- 通过调节教师模型的隐单元数和权重行相关性,控制数据结构复杂度。
- 数据量需随教师模式数与相关性增加而提升,低温推理会阻碍模式学习。
- 支持一对一或一对多模式学习,适用于验证彩票理论等机制研究。
受限玻尔兹曼机(RBM)是能够学习具有丰富内在结构数据的生成模型。本文研究教师-学生框架下,学生RBM从教师RBM生成的结构化数据中学习的问题。通过调整教师模型的隐单元数及权重行间的相关性(即模式),可调控数据的结构复杂度。当无相关性时,验证了性能与教师模式数、学生隐单元数无关的猜想,并认为该设置可作为研究彩票理论的简化模型。在有相关性的条件下,发现学习教师模式所需的临界数据量随模式数与相关性增加而减少。在两种情形下,即使数据量较大,若推理温度过低,仍无法学习教师模式。本框架支持学生模型以一对一或一对多方式学习教师模式,将此前针对两个隐单元的研究推广至任意有限数量。
原文摘要 · Abstract (English)
Restricted Boltzmann machines (RBM) are generative models capable to learn data with a rich underlying structure. We study the teacher-student setting where a student RBM learns structured data generated by a teacher RBM. The amount of structure in the data is controlled by adjusting the number of hidden units of the teacher and the correlations in the rows of the weights, a.k.a. patterns. In the absence of correlations, we validate the conjecture that the performance is independent of the number of teacher patters and hidden units of the student RBMs, and we argue that the teacher-student setting can be used as a toy model for studying the lottery ticket hypothesis. Beyond this regime, we find that the critical amount of data required to learn the teacher patterns decreases with both their number and correlations. In both regimes, we find that, even with a relatively large dataset, it becomes impossible to learn the teacher patterns if the inference temperature used for regularization is kept too low. In our framework, the student can learn teacher patterns one-to-one or many-to-one, generalizing previous findings about the teacher-student setting with two hidden units to any arbitrary finite number of hidden units.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。