证明大模型可压缩至极小规模,仍保持学习效果。
A universal compression theory for lottery ticket hypothesis and neural scaling laws
- 提出通用压缩理论,将大模型压缩至对数级宽度
- 压缩后模型学习动态不变,数据量可降至极低
- 为彩票票券假说提供理论支持,适合研究模型压缩者
大规模模型训练中,性能通常随参数量和数据集规模呈缓慢幂律增长。一个基础且关键的问题是:能否用显著更小的模型和更少的数据实现相当的性能?本文给出肯定且构造性的回答。我们证明:任意关于 $d$ 个对象的置换不变函数,可渐近压缩为仅含 $ ext{polylog} hinspace d$ 个对象的函数,且误差趋近于零,此压缩率被证明为最优。该定理带来两个核心推论:(Ia) 大型神经网络可压缩至多项式对数级宽度,同时保持其学习动态;(Ib) 大型数据集可压缩至多项式对数级规模,且对应模型的损失曲面不变。推论 (Ia) 直接证明了动态彩票票券假说,即任意普通网络均可强压缩,学习过程与结果不变;推论 (Ib) 表明神经缩放律 $L hicksim d^{-α}$ 可被提升为任意快的幂律衰减,最终达到 $ ext{exp}(-α' hinspace ext{root}[m]{d})$。
原文摘要 · Abstract (English)
When training large-scale models, the performance typically scales with the number of parameters and the dataset size according to a slow power law. A fundamental theoretical and practical question is whether comparable performance can be achieved with significantly smaller models and substantially less data. In this work, we provide a positive and constructive answer. We prove that a generic permutation-invariant function of $d$ objects can be asymptotically compressed into a function of $\operatorname{polylog} d$ objects with vanishing error, which is proved to be the optimal compression rate. This theorem yields two key implications: (Ia) a large neural network can be compressed to polylogarithmic width while preserving its learning dynamics; (Ib) a large dataset can be compressed to polylogarithmic size while leaving the loss landscape of the corresponding model unchanged. Implication (Ia) directly establishes a proof of the dynamical lottery ticket hypothesis, which states that any ordinary network can be strongly compressed such that the learning dynamics and result remain unchanged. (Ib) shows that a neural scaling law of the form $L\sim d^{-α}$ can be boosted to an arbitrarily fast power law decay, and ultimately to $\exp(-α' \sqrt[m]{d})$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。