Soro是专为塔吉克语设计的轻量级大模型,适合在低算力环境下部署。
Soro: A Lightweight Foundation Model and Chatbot for Tajik

- 基于Gemma 3进行塔吉克语持续预训练,使用19亿词元语料
- 在塔吉克语评测集上显著优于同尺寸基线模型,英文性能也保持良好
- 支持FP8/INT4量化,便于边缘设备部署,已用于教育试点
我们提出Soro,一系列专为塔吉克语设计的对话式大语言模型,旨在应对塔吉克斯坦算力与网络资源有限的现实挑战。基于开放权重的Gemma 3检查点,我们在19亿词元的塔吉克语语料库(包含筛选后的网页文本、PDF文档及课程对齐教育材料)上进行仅塔吉克语的持续预训练,随后在4万条塔吉克教师风格示例上进行监督指令微调。为克服标准基准中塔吉克语覆盖不足的问题,我们构建了一套涵盖通用知识、语言能力及中小学和大学入学考试领域的塔吉克语评测集,并在Hugging Face上开源。在这些塔吉克语评测集上,Soro显著优于同规模的Gemma 3基线模型,同时在标准英语数据集上仍保持较强性能。我们进一步证明,对Soro采用FP8和INT4量化可在保留大部分塔吉克语性能的同时显著降低内存占用,支持当前教育领域试点项目,并计划在塔吉克斯坦全国学校推广。
原文摘要 · Abstract (English)
We present Soro, a family of Tajik-specialized conversational large language models (LLMs) designed for real-world deployment under tight compute and connectivity constraints in Tajikistan. Starting from open-weight Gemma 3 checkpoints, we perform Tajik-only continual pretraining on a curated 1.9-billion-token corpus spanning filtered web text, PDF documents, and curriculum-aligned educational materials, followed by supervised instruction tuning on 40K Tajik teacher-style examples. To enable rigorous evaluation despite the limited coverage of Tajik in standard benchmarks, we introduce a suite of Tajik benchmarks covering general knowledge, linguistic competence, and school- and university entrance-exam domains, and we open-source them on Hugging Face. Across these Tajik benchmarks, Soro substantially outperforms same-size Gemma 3 baselines while retaining strong English performance on standard datasets. We further show that FP8 and INT4 quantization of Soro preserves most Tajik-language gains while reducing memory requirements for edge deployment, supporting an ongoing education-sector pilot and planned scale-out across schools in Tajikistan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。