arXiv:2606.01302cs.LG2026-06

发现小模型性能与内部结构随规模变化的关联性。

Structure and Scale in Simplicial Sequence Modelling

论文配图:Structure and Scale in Simplicial Sequence Modelling
图 1 · 摘自论文原文
  • 用小型Transformer预测隐马尔可夫模型输出,研究规模扩展时行为与结构关系。
  • 观察到性能提升与残差激活在概率单纯形中的线性编码模式一致。
  • 为理解大模型的可预测行为提供结构层面的新视角,适合对机制解释感兴趣的研究者。

现代大规模深度学习表现出两个显著的实证现象:行为缩放定律(规模增大时性能可预测地提升)和涌现机制(深度神经网络中出现有结构的内部表征与计算回路)。我们假设这两个现象是相关的:行为的可预测变化源于内部计算结构的可预测变化。本文报告了这一假设的初步证据:在训练小型Transformer以预测隐藏马尔可夫模型输出时,发现性能的缩放模式与残差激活在概率单纯形中线性编码潜在状态信念分布的现象存在相关性。

原文摘要 · Abstract (English)

Modern large-scale deep learning exhibits two striking empirical phenomena: behavioural scaling laws (predictable performance gains with increasing scale) and emergent mechanisms (structured internal representations and circuits in deep neural networks). We hypothesise that these two phenomena are connected: that predictable changes in behaviour are the result of predictable changes in internal computational structure. In this paper, we report preliminary evidence of such a connection. We find a correlation between scaling patterns in performance and representations in small transformers trained to predict the outputs of a hidden Markov model, for which residual activations are known to linearly encode a belief distribution over latent states in a probability simplex.

Transformer缩放定律结构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。