arXiv:2505.02410cs.CLcs.AI2025-05被引 2

110亿参数波兰语模型,性能超越更大模型。

Bielik 11B v2 Technical Report

  • 基于Mistral架构扩展至110亿参数,采用加权指令损失与自适应学习率优化训练。
  • 在波兰语多项评测中表现优异,优于多款参数量2-6倍的模型。
  • 支持多种量化方案,适合资源受限设备部署,助力小语种AI发展。

我们提出Bielik 11B v2,一款针对波兰语文本处理的先进语言模型。基于Mistral 7B v0.2架构,通过深度扩展提升至110亿参数,该模型在波兰语基准测试中表现卓越,同时具备强大的跨语言能力。我们引入两项关键技术:加权指令交叉熵损失,根据训练样本质量动态赋权;自适应学习率,随上下文长度动态调整。在多个基准上的综合评估显示,Bielik 11B v2性能超越多款参数量达其2-6倍的模型,并在从语言理解到复杂推理的任务中显著优于其他专用波兰语模型。模型具备参数高效性及广泛的量化选项,可适配不同硬件配置,推动波兰语AI能力发展,为低资源语言的高效建模树立新标杆。

原文摘要 · Abstract (English)

We present Bielik 11B v2, a state-of-the-art language model optimized for Polish text processing. Built on the Mistral 7B v0.2 architecture and scaled to 11B parameters using depth up-scaling, this model demonstrates exceptional performance across Polish language benchmarks while maintaining strong cross-lingual capabilities. We introduce two key technical innovations: Weighted Instruction Cross-Entropy Loss, which optimizes learning across diverse instruction types by assigning quality-based weights to training examples, and Adaptive Learning Rate, which dynamically adjusts based on context length. Comprehensive evaluation across multiple benchmarks demonstrates that Bielik 11B v2 outperforms many larger models, including those with 2-6 times more parameters, and significantly surpasses other specialized Polish language models on tasks ranging from linguistic understanding to complex reasoning. The model's parameter efficiency and extensive quantization options enable deployment across various hardware configurations, advancing Polish language AI capabilities and establishing new benchmarks for resource-efficient language modeling in less-represented languages.

波兰语大模型参数效率量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。