arXiv:2505.07070cs.LGcond-mat.dis-nn2025-05被引 5

对比卷积与注意力模型在层级语言中的学习效率,发现结构匹配者更快收敛。

Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures

  • 用随机层级模型生成数据,理论推导性能随规模变化规律
  • 卷积网络因局部性与权共享,性能提升速度优于变压器
  • 揭示模型架构与数据统计特性如何共同决定学习效率

神经语言模型在预测下一个词时如何习得语言结构?我们通过随机层级模型(RHM)——一种能捕捉自然语言层级结构但又保持解析可处理性的概率上下文无关语法集合——生成合成数据,并推导出神经网络性能的理论缩放定律。此前我们提出了基于数据相关性的表征学习理论,解释深度学习模型如何逐层顺序捕获数据的层级结构。本文将该理论框架扩展至考虑架构差异:预测并实证验证了卷积网络因结构与生成过程的局部性和权共享相契合,其性能缩放速度优于依赖全局自注意力机制的变压器模型。这一发现阐明了神经缩放定律背后的架构偏置,凸显了模型架构与数据统计特性之间的相互作用对表征学习的影响。

原文摘要 · Abstract (English)

How do neural language models acquire a language's structure when trained for next-token prediction? We address this question by deriving theoretical scaling laws for neural network performance on synthetic datasets generated by the Random Hierarchy Model (RHM) -- an ensemble of probabilistic context-free grammars designed to capture the hierarchical structure of natural language while remaining analytically tractable. Previously, we developed a theory of representation learning based on data correlations that explains how deep learning models capture the hierarchical structure of the data sequentially, one layer at a time. Here, we extend our theoretical framework to account for architectural differences. In particular, we predict and empirically validate that convolutional networks, whose structure aligns with that of the generative process through locality and weight sharing, enjoy a faster scaling of performance compared to transformer models, which rely on global self-attention mechanisms. This finding clarifies the architectural biases underlying neural scaling laws and highlights how representation learning is shaped by the interaction between model architecture and the statistical properties of data.

层级结构缩放定律模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。