arXiv:2605.27989cs.LG2026-05

模型深度宽度比影响资源利用效率与泛化能力

Law of Neural Interaction: Depth-Width Shape, Interaction Efficiency, and Generalization

论文配图:Law of Neural Interaction: Depth-Width Shape, Interaction Efficiency, and Generalization
图 1 · 摘自论文原文
  • 提出神经交互概念,将超叠加从参数空间扩展到梯度空间
  • 固定预算下,最优深度宽度比使模型实现高效交互与更好泛化
  • 该比例在大预算下保持稳定,适合指导模型架构设计

规模定律指导下的现代大语言模型资源消耗日益增加,但其在固定预算下是否高效利用资源仍存疑。已有研究证明超叠加是损失的关键因素。我们基于神经特征假设,将超叠加从参数空间拓展至梯度空间,定义为神经交互。发现固定预算下,良好泛化通常伴随高效神经交互,通过调节深度-宽度比($R_{D/W}$)可使模型进入高效交互区间。此外,随着预算扩大,该高效区间相对稳定。对比现有小规模稠密LLM发现,位于此区间的模型在MMLU-Pro基准上表现更优。结果表明,$R_{D/W}$影响资源利用效率,进而决定泛化性能,为模型形状初始化及泛化机制理解提供新视角。代码已公开于:https://anonymous.4open.science/r/Neural_Interaction_Law-D788

原文摘要 · Abstract (English)

The guidance of scaling laws has increased the resource demands of modern large language models (LLMs), yet it remains questionable whether these models utilize resources effectively under a fixed budget. Previous research has proved superposition as a key contributor to loss. By leveraging the Neural Feature Ansatz, we extend superposition from parameter space to gradient space and define it as neural interaction. We find that under a fixed budget, good generalization is usually accompanied by efficient neural interactions, and the model can be placed in an efficient interaction interval by adjusting its depth-width ratio ($R_{D/W}$). In addition, as the budget scales up, the efficient interaction interval of the model remains relatively stable. By comparing existing small scale dense LLMs, we observe that models operating near this interval tend to perform better on the MMLU-Pro benchmark. Our findings reveal that the $R_{D/W}$ influences resource utilization efficiency and thereby affects generalization, providing insights into model shape initialization and the understanding of model generalization mechanisms. Code for Neural Interaction Law is available at: https://anonymous.4open.science/r/Neural_Interaction_Law-D788

大模型架构设计泛化能力深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。