arXiv:2505.07861cs.CLcs.AI2025-05被引 7

用极小代价恢复高效推理中丢失的数学推理能力

Scalable LLM Reasoning Acceleration with Low-rank Distillation

  • 通过低秩蒸馏在前馈层恢复被压缩模型丢失的推理能力
  • 仅用20K合成数据和1%额外参数,就几乎完全恢复数学性能
  • 适合追求高效推理且不希望牺牲数学能力的研究者

由于生成过程长,大语言模型(LLM)在数学推理任务上需要大量计算资源和时间。尽管已有多种高效推理方法在语言任务上表现良好,但通常会严重损害数学推理能力。本文提出Caprese,一种资源高效的蒸馏方法,专注于前馈块的修复。在不改变原始权重、仅增加约1%参数、使用20K合成训练样本的情况下,能够恢复因高效推理部署而损失的大部分甚至全部推理能力,同时对指令型LLM的语言任务无负面影响。此外,Caprese将活跃参数数量大幅减少(如Gemma 2 9B和Llama 3.1 8B减少约20亿),可无缝集成到现有模型层中,降低延迟(next-token时间减少超16%),并促进输出更简洁(最多减少8.5%的词元数)。

原文摘要 · Abstract (English)

Due to long generations, large language model (LLM) math reasoning demands significant computational resources and time. While many existing efficient inference methods have been developed with excellent performance preservation on language tasks, they often severely degrade math performance. In this paper, we propose Caprese, a resource-efficient distillation method to recover lost capabilities from deploying efficient inference methods, focused primarily in feedforward blocks. With original weights unperturbed, roughly 1% of additional parameters, and only 20K synthetic training samples, we are able to recover much if not all of the reasoning capabilities lost from efficient inference for thinking LLMs and without harm to language tasks for instruct LLMs. Moreover, Caprese slashes the number of active parameters (~2B cut for Gemma 2 9B and Llama 3.1 8B) and integrates cleanly into existing model layers to reduce latency (>16% time-to-next-token reduction) while encouraging response brevity (up to 8.5% fewer tokens).

推理加速低秩蒸馏数学推理模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。