揭示分层生成数据下学习曲线的幂律规律,统一神经网络缩放定律理论。
Learning curves theory for hierarchically compositional data with power-law distributed features
- 基于概率上下文无关语法建模分层生成数据,分析学习曲线行为。
- 分类任务中学习曲线幂律指数由规则分布决定,常数受层级结构影响。
- 适用于语言、图像等具有分层结构的数据,适合研究模型缩放机制的读者。
近期理论认为,当任务可线性分解为幂律分布的单元时,会涌现出神经网络缩放定律。另一种情况是数据具有分层组合结构,如语言和图像中所见。为统一这两种观点,本文考虑基于概率上下文无关语法(PCFG)的分类与下一词预测任务——该语法通过层级生成规则生成数据。对于分类任务,我们证明:当生成规则呈幂律分布时,学习曲线呈现幂律形式,其指数取决于规则分布,而大尺度乘性常数则由层级结构决定。相比之下,在下一词预测任务中,规则分布仅影响学习曲线的局部细节,不决定大尺度行为的指数。
原文摘要 · Abstract (English)
Recent theories suggest that Neural Scaling Laws arise whenever the task is linearly decomposed into power-law distributed units. Alternatively, scaling laws also emerge when data exhibit a hierarchically compositional structure, as is thought to occur in language and images. To unify these views, we consider classification and next-token prediction tasks based on probabilistic context-free grammars -- probabilistic models that generate data via a hierarchy of production rules. For classification, we show that having power-law distributed production rules results in a power-law learning curve with an exponent depending on the rules' distribution and a large multiplicative constant that depends on the hierarchical structure. By contrast, for next-token prediction, the distribution of production rules controls the local details of the learning curve, but not the exponent describing the large-scale behaviour.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。