arXiv:2505.02550cs.LGcs.AI2025-05被引 5

小模型也能高效处理波兰语,性能媲美大模型。

Bielik v3 Small: Technical Report

  • 用定制分词器和自适应学习率优化小模型训练
  • 1.5B和4.5B模型在多项波兰语基准上表现接近大模型
  • 适合资源有限但需高质量波兰语AI的场景

我们提出Bielik v3系列参数高效生成式文本模型(1.5B和4.5B),专为波兰语处理优化。这些模型在仅2920亿个标记(来自3.03亿份文档)的精心构建语料上训练,通过自定义波兰语分词器APT4提升分词效率,采用加权指令交叉熵损失平衡不同指令类型的学习,并引入基于训练进度自适应调整的学习率。4.5B模型性能可与规模达其2-3倍的模型比肩,1.5B模型虽极小仍表现强劲。在Open PL LLM Leaderboard、复杂波兰语理解基准、Polish EQ-Bench及波兰医疗基准上均取得优异结果。该工作为低资源语言的高效语言建模树立新标准,使高精度波兰语AI更适用于计算资源受限的应用。

原文摘要 · Abstract (English)

We introduce Bielik v3, a series of parameter-efficient generative text models (1.5B and 4.5B) optimized for Polish language processing. These models demonstrate that smaller, well-optimized architectures can achieve performance comparable to much larger counterparts while requiring substantially fewer computational resources. Our approach incorporates several key innovations: a custom Polish tokenizer (APT4) that significantly improves token efficiency, Weighted Instruction Cross-Entropy Loss to balance learning across instruction types, and Adaptive Learning Rate that dynamically adjusts based on training progress. Trained on a meticulously curated corpus of 292 billion tokens spanning 303 million documents, these models excel across multiple benchmarks, including the Open PL LLM Leaderboard, Complex Polish Text Understanding Benchmark, Polish EQ-Bench, and Polish Medical Leaderboard. The 4.5B parameter model achieves results competitive with models 2-3 times its size, while the 1.5B model delivers strong performance despite its extremely compact profile. These advances establish new benchmarks for parameter-efficient language modeling in less-represented languages, making high-quality Polish language AI more accessible for resource-constrained applications.

波兰语小模型参数效率生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。