arXiv:2603.11881cs.CLcs.AI2026-03被引 3

将波兰语大模型压缩33.4%,性能损失仅10%且提速50%

Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language

  • 分两阶段压缩:结构化剪枝+基于日志的蒸馏
  • 参数从11.04B减至7.35B,恢复90%原始性能
  • 适合资源有限场景下的小语种模型部署

本报告介绍Bieli k-Minitron-7B的构建过程,这是针对欧洲语言优化的Bielik-11B-v3.0模型的压缩版本,参数量为7.35B。通过借鉴NVIDIA Minitron方法的两阶段压缩策略,结合结构化混合剪枝与知识蒸馏,将参数量减少33.4%(从11.04B降至7.35B)。我们使用NVIDIA Model Optimizer进行结构剪枝,采用NVIDIA NeMo框架实现基于日志的蒸馏以恢复性能。蒸馏后,模型经过监督微调(SFT)、直接偏好优化(DPO-P)和强化学习(GRPO)的对齐流程。最终模型在保持约90%基线性能的同时,实现最高达50%的推理加速。该方法为低资源语言模型提供了高效、高质量的部署路径。

原文摘要 · Abstract (English)

This report details the creation of Bielik-Minitron-7B, a compressed 7.35B parameter version of the Bielik-11B-v3.0 model, specifically optimized for European languages. By leveraging a two-stage compression methodology inspired by the NVIDIA Minitron approach, we combined structured hybrid pruning and knowledge distillation to reduce the model's parameter count by 33.4%, from 11.04B to 7.35B. We utilized the NVIDIA Model Optimizer for structural pruning and the NVIDIA NeMo Framework for logit-based distillation for quality recovery. Following distillation, the model underwent a rigorous alignment pipeline consisting of Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO-P), and Reinforcement Learning (GRPO). Our final model successfully recovered approximately 90% of the baseline model's performance while providing up to 50% inference speedup. This approach demonstrates an efficient pathway to create language models for less-represented languages, preserving the original model quality while reducing inference deployment costs.

模型压缩波兰语知识蒸馏推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。