arXiv:2507.02982cs.CL2025-07

用知识蒸馏压缩大模型,让小模型高效解数学应用题

We Need Knowledge Distillation for Solving Math Word Problems

  • 用BERT编码向量做知识蒸馏,训练小型学生模型
  • 小模型仅需1/12参数,性能达教师模型的90%
  • 方法通用性强,适合教育类智能辅导系统落地

大型语言模型(LLMs)在数学能力上的提升推动了中小学智能辅导系统的发展。然而,其高昂的计算成本限制了教育场景的应用。本文研究通过知识蒸馏压缩解决数学应用题(MWPs)的大型模型。我们对BERT生成的嵌入向量进行压缩,并蒸馏出一个更小的学生模型。结果表明,该学生模型仅使用教师模型1/12的参数,即可保持近90%的性能表现。此外,压缩后的向量在各类与数学应用题相关的任务中均表现出良好泛化能力,且蒸馏过程不依赖特定任务。这说明所用原理具有通用性。我们进一步分析发现,词性信息而非实体识别是影响可压缩性的关键因素,这可能解释了嵌入向量的可压缩性。效率提升与成本降低为智能教育系统带来显著价值。

原文摘要 · Abstract (English)

The enhancement of mathematical capabilities in large language models (LLMs) fosters new developments in mathematics education within primary and secondary schools, particularly as they relate to intelligent tutoring systems. However, LLMs require substantial computational resources, resulting in significant costs in educational contexts. To mitigate this drawback, this paper investigates the feasibility of compressing LLMs for solving math word problems (MWPs). We compress the embedded vectors encoded by BERT and distill a considerably smaller student model. Our findings indicate that the student model can maintain nearly 90% of the performance of the teacher model while utilizing only 1/12 of its parameters. In addition to achieving high accuracy, the model exhibits strong generalizability, as the compressed vectors perform well across all tasks related to MWPs, and the distillation process is not task-specific. The success of this distillation demonstrates that the underlying principles are generic and not limited to a specific task. We further explore the reasons behind the compressibility of embedded vectors, revealing that part-of-speech information, rather than entity recognition, is crucial for MWPs, which may significantly contribute to their compressibility. The improvements in efficiency and cost reduction provide substantial value for intelligent tutoring systems and significantly advance the field of intelligent education.

知识蒸馏数学推理智能教育模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。