提出统一压缩表示的缩放定律,用容量指标预测模型效率。
Unified Scaling Laws for Compressed Representations
- 用容量指标衡量压缩表示的拟合能力,统一预测性能
- 验证多种压缩格式(稀疏、量化等)下缩放定律可组合适用
- 适用于模型压缩算法优化,尤其适合追求高效推理的研究者
缩放定律推动了机器学习的发展,使模型性能可基于模型规模、计算量和数据量进行预测。随着人工智能计算成本上升,量化与稀疏化等模型压缩技术应运而生,以缓解大规模训练与推理的高算力需求。本文研究缩放定律与压缩表示之间的关系,探讨是否能建立统一框架,准确预测在稀疏、标量量化、稀疏量化甚至向量量化等不同压缩表示下的模型表现。核心贡献包括验证通用缩放定律形式,并证明其在各类压缩方式中独立及可组合适用。主要发现是:理论上与实证上均表明,一个基于拟合随机高斯数据能力的简单“容量”指标,可稳健预测多种压缩表示下的参数效率。此外,我们扩展该框架,直接比较不同压缩格式的精度潜力,并推导出更优的稀疏量化训练算法。
原文摘要 · Abstract (English)
Scaling laws have shaped recent advances in machine learning by enabling predictable scaling of model performance based on model size, computation, and data volume. Concurrently, the rise in computational cost for AI has motivated model compression techniques, notably quantization and sparsification, which have emerged to mitigate the steep computational demands associated with large-scale training and inference. This paper investigates the interplay between scaling laws and compression formats, exploring whether a unified scaling framework can accurately predict model performance when training occurs over various compressed representations, such as sparse, scalar-quantized, sparse-quantized or even vector-quantized formats. Our key contributions include validating a general scaling law formulation and showing that it is applicable both individually but also composably across compression types. Based on this, our main finding is demonstrating both theoretically and empirically that there exists a simple "capacity" metric -- based on the representation's ability to fit random Gaussian data -- which can robustly predict parameter efficiency across multiple compressed representations. On the practical side, we extend our formulation to directly compare the accuracy potential of different compressed formats, and to derive better algorithms for training over sparse-quantized formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。