arXiv:2507.19699cs.CL2025-07被引 3

评估压缩模型在多语言环境中的表现,发现多语言模型更优,量化有效但剪枝损害性能。

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks

  • 对比多语言与单语言模型在阿拉伯语、印地语等语言上的表现。
  • 4/8位量化保持精度,但激进剪枝显著降低大模型性能。
  • 适合关注公平性与低资源语言NLP的研究者参考。

尽管大语言模型在高资源语言上取得显著进展,其在低资源语言如卡纳达语和阿拉伯语中的能力仍不明确。本研究在阿拉伯语、英语和印地语等语言上基准测试了多语言与单语言大模型(如BLOOMZ、AceGPT、Jais、LLaMA-2、XGLM、AraGPT2)的性能,重点分析剪枝与量化等模型压缩策略的影响。结果表明,语言多样性和资源可用性显著影响SOTA大模型的表现。多语言模型整体优于语言专用模型,体现显著的跨语言迁移优势。4位与8位量化能有效维持精度并提升效率,但激进剪枝会严重损害性能,尤其在大模型中更为明显。研究指明构建可扩展且公平的多语言NLP解决方案的关键策略,并强调需针对低资源场景中的幻觉与泛化误差采取干预措施。

原文摘要 · Abstract (English)

Although LLMs have attained significant success in high-resource languages, their capacity in low-resource linguistic environments like Kannada and Arabic is not yet fully understood. This work benchmarking the performance of multilingual and monolingual Large Language Models (LLMs) across Arabic, English, and Indic languages, with particular emphasis on the effects of model compression strategies such as pruning and quantization. Findings shows significant performance differences driven by linguistic diversity and resource availability on SOTA LLMS as BLOOMZ, AceGPT, Jais, LLaMA-2, XGLM, and AraGPT2. We find that multilingual versions of the model outperform their language-specific counterparts across the board, indicating substantial cross-lingual transfer benefits. Quantization (4-bit and 8-bit) is effective in maintaining model accuracy while promoting efficiency, but aggressive pruning significantly compromises performance, especially in bigger models. Our findings pinpoint key strategies to construct scalable and fair multilingual NLP solutions and underscore the need for interventions to address hallucination and generalization errors in the low-resource setting.

多语言NLP模型压缩低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。