发现大模型前馈层潜空间利用率存在不对称规律,指导高效设计。
Spectral Scaling Laws in Language Models: How Effectively Do Feed-Forward Networks Use Their Latent Space?
- 用谱分析法量化前馈网络激活的潜在方向数量。
- 软秩随宽度近似幂律增长,硬秩却增速缓慢且波动大。
- 揭示主流模型广泛浪费潜空间,适合模型压缩与结构优化研究者。
随着大语言模型规模扩大,问题不仅在于模型有多大,更在于其容量是否被有效利用。现有缩放定律仅关联模型大小与损失,忽视各组件对潜空间的实际利用情况。本文聚焦前馈网络(FFNs),将宽度选择重构为谱利用问题。通过轻量诊断工具——硬秩(参与度比率)、软秩(香农秩)、谱集中度及综合谱利用指数(SUI),我们量化了LLaMA、GPT-2和nGPT系列模型中被有意义激活的潜空间方向数。核心发现是一种非对称谱缩放规律:软秩随FFN宽度近乎完美遵循幂律,而硬秩仅亚线性增长且方差显著。这表明加宽FFN主要增加低能量尾部方向,主导模式子空间早期即饱和。此外,在更大宽度下,方差进一步坍缩至狭窄子空间,导致大量潜空间未被利用。该结果将FFN宽度选择重新定义为尾部容量与主导模式容量之间的权衡,为推理高效的大语言模型设计提供明确指导。
原文摘要 · Abstract (English)
As large language models (LLMs) scale, the question is not only how large they become, but how much of their capacity is effectively utilized. Existing scaling laws relate model size to loss, yet overlook how components exploit their latent space. We study feed-forward networks (FFNs) and recast width selection as a spectral utilization problem. Using a lightweight diagnostic suite -- Hard Rank (participation ratio), Soft Rank (Shannon rank), Spectral Concentration, and the composite Spectral Utilization Index (SUI) -- we quantify how many latent directions are meaningfully activated across LLaMA, GPT-2, and nGPT families. Our key finding is an asymmetric spectral scaling law: soft rank follows an almost perfect power law with FFN width, while hard rank grows only sublinearly and with high variance. This asymmetry suggests that widening FFNs mostly adds low-energy tail directions, while dominant-mode subspaces saturate early. Moreover, at larger widths, variance further collapses into a narrow subspace, leaving much of the latent space under-utilized. These results recast FFN width selection as a principled trade-off between tail capacity and dominant-mode capacity, offering concrete guidance for inference-efficient LLM design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。