arXiv:2603.04162cs.CLcs.AI2026-03

首次系统评估波兰语大模型2比特量化,性能接近原模型且部分方法更省空间。

Bielik-Q2-Sharp: A Comparative Study of Extreme 2-bit Quantization Methods for a Polish 11B Language Model

  • 对比6种2比特量化方法,基于波兰语数据集和共享海森矩阵进行评估。
  • 最佳方案在22项基准上达71.92%,推理能力提升3.6个百分点,仅增加0.66GB内存。
  • 发现旋转类方法保对数似然但生成质量崩溃,适合关注压缩效率的研究者。

我们提出Bielik-Q2-Sharp,首次系统性评估极端2比特量化在波兰语大语言模型上的应用。以Bielik-11B-v2.3-Instruct(11B参数,Mistral架构)为基线模型,比较六种先进后训练量化方法——QuIP#、SpinQuant+GPTQ、ButterflyQuant、QTIP、VPTQ和AQLM——均在波兰语语料库CulturaX-PL上校准,并使用共享海森矩阵。最优变体QuIP# E8P12在22个波兰语基准上达到71.92%准确率,仅略低于IQ2_XXS基线的72.07%(统计噪声范围内),体积增加3.26 GB(基线约2.6 GB)。在eq_bench上得分47.14,优于基线43.53(+3.6pp),表明更高阶推理能力更好保留。QTIP实现最高每比特效率(~2.4 bpw时79.4% MC acc_norm,3.27 GB),性能媲美VPTQ但体积小35%。此外,发现旋转类方法存在生成与概率质量解耦现象:保对数似然但自回归生成严重失败。整个研究由单人独立完成,使用vast.ai云GPU,在285美元预算内完成。所有模型、海森矩阵及评估日志均已公开。

原文摘要 · Abstract (English)

We present Bielik-Q2-Sharp, the first systematic academic evaluation of extreme 2-bit quantization applied to a Polish large language model. Using Bielik-11B-v2.3-Instruct (11B parameters, Mistral architecture) as our base model, we compare six state-of-the-art post-training quantization methods -- QuIP#, SpinQuant+GPTQ, ButterflyQuant, QTIP, VPTQ, and AQLM -- all calibrated on a Polish-language corpus (CulturaX-PL) with shared Hessian matrices. Our best variant (QuIP# E8P12) achieves 71.92% across 22 Polish benchmarks versus 72.07% for the IQ2_XXS baseline -- within statistical noise, at a modest size premium (3.26 GB vs. ~2.6 GB). On eq_bench, our method scores 47.14 versus 43.53 (+3.6pp), suggesting superior preservation of higher-order reasoning. QTIP achieves the best per-bit efficiency (79.4% MC acc_norm at ~2.4 bpw, 3.27 GB), matching VPTQ's quality at 35% smaller size. We additionally document a MC-generation dissociation phenomenon where rotation-based methods preserve log-likelihood quality but fail catastrophically at autoregressive generation. The entire project was conducted by a single independent researcher on cloud GPUs (vast.ai) within a $285 budget. All models, Hessians, and evaluation logs are publicly available.

2比特量化波兰语模型高效推理模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。