arXiv:2502.18424cs.CL2025-02Conference of the …被引 3

提出新校准方法MixCal,让压缩后的语言模型在专业领域表现更好

Compressing Language Models for Specialized Domains

  • 提出后训练校准方法MixCal,无需全参数微调
  • 在生物医学等专业任务上显著提升压缩模型性能
  • 既提升领域性能又降低压缩计算成本,适合部署场景

语言模型在多领域表现出色,但推理时需大量计算资源。剪枝和量化等压缩技术虽能提升效率,但在生物医学、法律等专业领域常导致性能下降。现有解决方案依赖昂贵的全参数微调。为此,本文提出MixCal,一种后训练校准方法,旨在提升压缩语言模型在特定领域的表现。大量实验表明,MixCal在领域任务上显著优于现有方法,同时保持通用性能。更重要的是,该方法还能降低压缩过程的计算开销。

原文摘要 · Abstract (English)

Language models (LMs) excel at tasks across diverse domains, yet require substantial computational resources during inference. Compression techniques such as pruning and quantization offer a practical path towards efficient LM deployment, exemplified by their ability to preserve performance on general-purpose benchmarks. However, general-purpose LM compression methods can negatively affect performance in specialized domains (e.g. biomedical or legal). Recent work has sought to address this issue, but requires a computationally expensive full-parameter fine-tuning pipeline. To this end, we propose MixCal, a novel calibration method designed to improve the in-domain performance of compressed LMs in a post-training setting. Through extensive experimentation, we demonstrate that MixCal substantially outperforms existing approaches on domain-specific tasks and preserves general performance. Notably, these performance gains are achieved while also reducing the computational cost of LM compression.

模型压缩领域适配后训练高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。