揭示了序列模型中数据模式与损失曲面几何的关系。
Modes of Sequence Models and Learning Coefficients
- 用希尔伯特空间和张量分解识别数据的主要模式
- 小振幅模式被截断后仍保留主要结构,有效分布更简洁
- 实际学习系数反映的是有效分布的几何特征
我们建立了一个序列建模的几何框架,将数据中的模式与变换器网络损失曲面的可测量特性联系起来。首先,将条件序列分布置于希尔伯特空间框架中,并应用张量分解以识别其主要模式;截断低振幅模式后得到一个保留主导结构但去除统计细节的有效数据分布。其次,理论上证明局部学习系数(LLC)估计对低于数据相关阈值的模式不敏感,因此实践中计算的LLC刻画的是有效而非真实分布的几何特征。这一洞察解释了为何即使网络参数并非严格最小化总体损失,仍可获得可靠的LLC估计,并揭示了SGLD中的逆温度在调节景观结构分辨率上的作用。
原文摘要 · Abstract (English)
We develop a geometric account of sequence modelling that links patterns in the data to measurable properties of the loss landscape in transformer networks. First, we cast conditional sequence distributions into a Hilbert-space framework and apply tensor decompositions to identify their principal modes. Truncating the small-amplitude modes yields an effective data distribution that preserves dominant structure while discarding statistical detail. Second, we show theoretically that Local Learning Coefficient (LLC) estimates are insensitive to modes below a data-dependent threshold. Consequently, the LLC calculated in practice characterises the geometry of the effective rather than the true distribution. This insight clarifies why reliable LLC estimates can be obtained even when a network parameter is not a strict minimiser of the population loss, and it highlights how the inverse temperature in SGLD acts as a resolution dial on the landscape structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。