arXiv:2601.01244cs.CL2026-01

用轻量持续预训练让匈牙利语大模型在普通硬件上跑起来

Racka: Efficient Hungarian LLM Adaptation on Academic Infrastructure

  • 基于Qwen-3 4B用低秩适配(LoRA)实现高效持续预训练
  • 1600亿子词训练,匈牙利语占比44%,保持英德语性能
  • 适合资源有限的学术机构做小语种大模型适配

我们提出Racka,一种轻量级、持续预训练的大语言模型,旨在缩小匈牙利语与英语、德语等高资源语言之间的差距。Racka采用低秩适配(LoRA)在Qwen-3 4B基础上进行参数高效持续预训练,可在配备A100(40GB)的高性能计算集群上运行,且对节点间带宽要求低。为更好匹配训练分布,我们替换并优化了分词器,显著提升匈牙利语的分词效率,同时保持英语和德语的竞争力。模型在1600亿子词上训练,数据来源包括互联网和高质量精选内容,其中匈牙利语占44%,英语24%,德语21%,代码11%。该数据构成有助于缓解灾难性遗忘,维持高资源语言能力。初步结果表明语言适配效果稳定但尚有提升空间。

原文摘要 · Abstract (English)

We present Racka, a lightweight, continually pretrained large language model designed to bridge the resource gap between Hungarian and high-resource languages such as English and German. Racka employs parameter-efficient continual pretraining via Low-Rank Adaptation (LoRA) on a Qwen-3 4B backbone, making the recipe practical on A100 (40GB)-based HPC clusters with low inter-node bandwidth. To better match the training distribution, we replace and adapt the tokenizer, achieving substantially improved tokenization fertility for Hungarian while maintaining competitive performance in English and German. The model is trained on 160B subword tokens drawn from a mixture of internet and high-quality curated sources, with a composition of 44% Hungarian, 24% English, 21% German, and 11% code. This data mix is chosen to mitigate catastrophic forgetting and preserve high-resource language capabilities during continual pretraining. Our preliminary results indicate modest but stable results in language adaptation.

大模型小语种持续预训练LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。