arXiv:2511.03823cs.CLcs.AI2025-11被引 10

波兰首个开源大模型家族,专为波兰语打造。

PLLuM: A Family of Polish Large Language Models

  • 构建1400亿词元波兰语语料库,支持大模型预训练
  • 推出7.7万条指令数据与10万条偏好优化数据集
  • 含安全过滤模块,适合政策、公共事务等本土应用

大型语言模型在现代人工智能中扮演核心角色,但其发展主要集中在英语,导致其他语言支持有限。我们提出PLLuM(波兰大语言模型),这是首个专为波兰语设计的开源基础模型家族。由波兰主要研究机构联合开发,旨在解决英语主导商业生态下高质量、透明且文化相关的语言模型需求。模型基于1400亿词元的波兰语语料库进行预训练,包含7.7万条定制指令数据和10万条偏好优化数据集。关键创新在于融合严格数据治理的负责任人工智能框架,以及输出修正与安全过滤的混合模块。本文详细阐述了基座模型与指令调优模型的架构、训练流程及对齐技术,并在公共行政下游任务中验证其有效性。通过公开发布,PLLuM致力于推动开放研究,强化波兰自主人工智能能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for other languages. We present PLLuM (Polish Large Language Model), the largest open-source family of foundation models tailored specifically for the Polish language. Developed by a consortium of major Polish research institutions, PLLuM addresses the need for high-quality, transparent, and culturally relevant language models beyond the English-centric commercial landscape. We describe the development process, including the construction of a new 140-billion-token Polish text corpus for pre-training, a 77k custom instructions dataset, and a 100k preference optimization dataset. A key component is a Responsible AI framework that incorporates strict data governance and a hybrid module for output correction and safety filtering. We detail the models' architecture, training procedures, and alignment techniques for both base and instruction-tuned variants, and demonstrate their utility in a downstream task within public administration. By releasing these models publicly, PLLuM aims to foster open research and strengthen sovereign AI technologies in Poland.

大模型波兰语开源AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。