arXiv:2604.10799cs.CLcs.AI2026-04

为波兰语优化分词器,提升大模型效率与效果

Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series

论文配图:Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
图 1 · 摘自论文原文
  • 采用专为波兰语设计的分词器替代通用分词方案
  • 7B和11B模型在波兰语任务上表现优于通用模型
  • 适合关注波兰语NLP或高效语言模型研究者

Bielik v3 PL系列(含7B和11B参数规模)是语言特化大模型优化的重要进展。通用大模型虽具多语言能力,但依赖通用分词器,难以捕捉波兰语的形态特征,导致词元化效率低、推理成本高、有效上下文窗口受限。本报告介绍了从基于Mistral的通用分词器转向专为波兰语优化的词汇表的过程,涵盖FOCUS嵌入初始化、多阶段预训练课程,以及后续通过监督微调、直接偏好优化和基于组相对策略优化的强化学习(带可验证奖励)进行对齐训练。

原文摘要 · Abstract (English)

The development of the Bielik v3 PL series, encompassing both the 7B and 11B parameter variants, represents a significant milestone in the field of language-specific large language model (LLM) optimization. While general-purpose models often demonstrate impressive multilingual capabilities, they frequently suffer from a fundamental architectural inefficiency: the use of universal tokenizers. These tokenizers, typically designed to cover a broad spectrum of languages, often fail to capture the morphological nuances of specific languages like Polish, leading to higher fertility ratios, increased inference costs, and restricted effective context windows. This report details the transition from the universal Mistral-based tokenization to a dedicated Polish-optimized vocabulary for the Bielik v3 models, exploring the FOCUS-based embedding initialization, the multi-stage pretraining curriculum, and the subsequent post-training alignment involving Supervised Fine-Tuning, Direct Preference Optimization, and Reinforcement Learning through Group Relative Policy Optimization with verifiable rewards.

语言模型波兰语分词器优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。