深度层在推理中不可或缺,但知识检索依赖浅层。
Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
- 按任务类型分析模型各层作用,发现不同任务依赖不同层次。
- 生成任务中,中深层对推理和长程连贯性至关重要。
- 知识检索靠浅层,推理能力可通过蒸馏重塑。
近期研究认为大语言模型的深层对表征学习贡献有限,可大幅剪枝而不损失性能。但此类结论多基于狭窄评估,可能忽略模型行为的关键方面。本文系统研究了模型深度利用在多种维度下的表现,涵盖评估协议、任务类别和模型架构。结果表明,深层通常不如浅层有效,但其贡献随评估设置显著变化:在仅基于似然的无生成评估中,剪枝大部分层仍可保持性能,仅有初始几层关键;而在生成评估中,中深层对推理能力和长程一致性起决定性作用。进一步发现,知识与检索集中于浅层,推理准确率则高度依赖深层,但可通过蒸馏重构。这些结果揭示模型深度使用具有高度异质性和情境依赖性,强调在解释与压缩大模型时需考虑任务、度量与模型自身特性。
原文摘要 · Abstract (English)
Recent studies suggest that the deeper layers of Large Language Models (LLMs) contribute little to representation learning and can often be removed without significant performance loss. However, such claims are typically drawn from narrow evaluations and may overlook important aspects of model behavior. In this work, we present a systematic study of depth utilization across diverse dimensions, including evaluation protocols, task categories, and model architectures. Our analysis confirms that very deep layers are generally less effective than earlier ones, but their contributions vary substantially with the evaluation setting. Under likelihood-based metrics without generation, pruning most layers preserves performance, with only the initial few being critical. By contrast, generation-based evaluation uncovers indispensable roles for middle and deeper layers in enabling reasoning and maintaining long-range coherence. We further find that knowledge and retrieval are concentrated in shallow components, whereas reasoning accuracy relies heavily on deeper layers -- yet can be reshaped through distillation. These results highlight that depth usage in LLMs is highly heterogeneous and context-dependent, underscoring the need for task-, metric-, and model-aware perspectives in both interpreting and compressing large models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。