让神经网络提前‘做梦’,应对数据变化,提升适应力。
Dreaming Learning
- 模拟潜在新数据分布,动态调整模型学习路径。
- 文本生成自相关提升约29%,模型收敛速度提高100%。
- 适合处理数据分布突变的场景,如时序变化、非平稳数据。
将新颖性引入深度学习系统仍具挑战。传统方法在非平稳数据源下易受干扰,难以保持模型稳定性。为此,我们提出受斯图尔特·考夫曼‘邻近可能’理论启发的训练算法——梦之学习(Dreaming Learning)。该方法在学习阶段探索新的数据空间,使神经网络能平滑接纳与预期统计特性不同的数据序列。其兼容的最大差异由探索阶段的采样温度参数决定。实验表明,该方法可有效应对马尔可夫链中的意外统计变化及文本序列的非平稳动态,在文本生成中自相关性能提升约29%,在马尔可夫范式突变情况下损失收敛速度提升约100%。
原文摘要 · Abstract (English)
Incorporating novelties into deep learning systems remains a challenging problem. Introducing new information to a machine learning system can interfere with previously stored data and potentially alter the global model paradigm, especially when dealing with non-stationary sources. In such cases, traditional approaches based on validation error minimization offer limited advantages. To address this, we propose a training algorithm inspired by Stuart Kauffman's notion of the Adjacent Possible. This novel training methodology explores new data spaces during the learning phase. It predisposes the neural network to smoothly accept and integrate data sequences with different statistical characteristics than expected. The maximum distance compatible with such inclusion depends on a specific parameter: the sampling temperature used in the explorative phase of the present method. This algorithm, called Dreaming Learning, anticipates potential regime shifts over time, enhancing the neural network's responsiveness to non-stationary events that alter statistical properties. To assess the advantages of this approach, we apply this methodology to unexpected statistical changes in Markov chains and non-stationary dynamics in textual sequences. We demonstrated its ability to improve the auto-correlation of generated textual sequences by $\sim 29\%$ and enhance the velocity of loss convergence by $\sim 100\%$ in the case of a paradigm shift in Markov chains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。