通过过参数化梯度下降,提升序列模型对信号结构的自适应能力。
Improving Adaptivity via Over-Parameterization in Sequence Models
- 在序列模型中引入过参数化梯度流,研究特征函数顺序的影响。
- 过参数化方法显著优于普通梯度流,且更深的过参数化提升泛化能力。
- 为理解神经网络的自适应性与泛化潜力提供新视角。
已知核函数的特征函数在核回归中起关键作用。我们通过多个例子表明,即使使用相同的特征函数集合,其排列顺序也会显著影响回归结果。通过将核矩阵对角化,我们提出一种在序列模型中应用的过参数化梯度下降方法,以捕捉固定特征函数集合下不同顺序的影响。理论分析显示,该过参数化梯度流能自适应地匹配信号内在结构,并显著优于标准梯度流。此外,更深层次的过参数化进一步增强了模型的泛化能力。这些结果不仅揭示了过参数化的潜在优势,还为理解神经网络在核外区域的自适应性与泛化潜力提供了新思路。
原文摘要 · Abstract (English)
It is well known that eigenfunctions of a kernel play a crucial role in kernel regression. Through several examples, we demonstrate that even with the same set of eigenfunctions, the order of these functions significantly impacts regression outcomes. Simplifying the model by diagonalizing the kernel, we introduce an over-parameterized gradient descent in the realm of sequence model to capture the effects of various orders of a fixed set of eigen-functions. This method is designed to explore the impact of varying eigenfunction orders. Our theoretical results show that the over-parameterization gradient flow can adapt to the underlying structure of the signal and significantly outperform the vanilla gradient flow method. Moreover, we also demonstrate that deeper over-parameterization can further enhance the generalization capability of the model. These results not only provide a new perspective on the benefits of over-parameterization and but also offer insights into the adaptivity and generalization potential of neural networks beyond the kernel regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。