arXiv:2601.11579cs.CLcs.AI2026-01被引 5

110亿参数模型专精波兰语,性能超越更大模型。

Bielik 11B v3: Multilingual Large Language Model for European Languages

  • 基于Mistral架构扩展至110亿参数,四阶段训练优化。
  • 波兰语任务表现超越同类模型,且优于2-6倍参数的大模型。
  • 支持多种量化,适合不同硬件部署,助力小语种AI发展。

我们提出Bielik 11B v3,一个针对波兰语高度优化的先进语言模型,同时保持对其他欧洲语言的强大能力。该模型在Mistral 7B v0.2架构基础上,通过深度扩展升级至110亿参数。开发过程包含连续预训练、监督微调(SFT)、直接偏好优化(DPO)和强化学习四个阶段。全面评估表明,Bielik 11B v3表现出色,在从基础语言理解到复杂推理的各类任务中显著超越其他专用波兰语模型,并优于许多参数量多出2至6倍的大型模型。其参数高效性结合丰富的量化选项,可实现跨多种硬件配置的有效部署。Bielik 11B v3不仅提升了波兰语的AI能力,还为资源较少的语言建立了高效高性能模型的新基准。

原文摘要 · Abstract (English)

We present Bielik 11B v3, a state-of-the-art language model highly optimized for the Polish language, while also maintaining strong capabilities in other European languages. This model extends the Mistral 7B v0.2 architecture, scaled to 11B parameters via depth up-scaling. Its development involved a comprehensive four-stage training pipeline: continuous pre-training, supervised fine-tuning (SFT), Direct Preference Optimization (DPO), and reinforcement learning. Comprehensive evaluations demonstrate that Bielik 11B v3 achieves exceptional performance. It significantly surpasses other specialized Polish language models and outperforms many larger models (with 2-6 times more parameters) on a wide range of tasks, from basic linguistic understanding to complex reasoning. The model's parameter efficiency, combined with extensive quantization options, allows for effective deployment across diverse hardware configurations. Bielik 11B v3 not only advances AI capabilities for the Polish language but also establishes a new benchmark for developing resource-efficient, high-performance models for less-represented languages.

大模型波兰语参数效率多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。