arXiv:2608.09941cs.CL2026-08

4bit量化让小模型在多语言下表现不一,低资源语言易崩溃。

The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs

  • 零样本评估8种语言,发现量化导致非拉丁语系表征崩溃
  • 低资源语言任务输出逻辑完全失效,精度损失超预期
  • 适合关注边缘设备多语言部署的开发者与研究者

尽管4比特权重量化对在边缘设备上部署小型语言模型至关重要,但现有性能退化评估仍高度集中于英语。我们对Gemma 4和Qwen 3.5架构进行了零样本多语言4比特量化评估,涵盖八种类型学差异显著的语言,使用MMLU ProX Lite和GlobalPIQA数据集。结果表明,参数截断暴露了预训练中的深层不平等。我们识别出四种现象:(1) 语言类型脆弱性:低资源及特定非拉丁文字在架构特异性双分离下发生表征坍缩,无法生成有效任务逻辑;(2) 母语脆弱性悖论:基础预训练路径未能提供足够精度保护;(3) 领域特异性遗忘:多步跨语言路由性能下降,而关联型社科类记忆保持稳健;(4) 量化抗性:高度饱和且类型对齐的领域抵抗确定性退化,量化后性能提升受统计噪声限制。

原文摘要 · Abstract (English)

While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a zero-shot multilingual evaluation of 4-bit quantization across the Gemma 4 and Qwen 3.5 architectures. Evaluating on eight typo-logically diverse languages using MMLU ProX Lite and GlobalPIQA, we show parameter truncation exposes deep pre-training inequalities. We identify four phenomena: (1) Typological Fragility: low-resource and specific non-Latin scripts suffer representational collapse via architecture-specific double dissociations, failing to generate valid task logits; (2) Home Language Fragility Paradox: foundational pre-training pathways provide limited precision loss protection; (3) Domain-Specific Forgetting: multi-step cross-lingual routing degrades while associative soft-science recall remains robust; and (4) Quantization Resistance: highly saturated, typologically aligned domains resist deterministic degradation, with post-quantization performance gains bounded by statistical noise.

量化多语言边缘计算模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。