首个开源亚美尼亚语大模型,用真实数据训练并公开全部配方。
From Zero to Hero: An Open LLM Ecosystem for Armenian
- 用437万篇新闻+37.3万道验证过的数理题微调基础模型
- 新模型在亚美尼亚语任务上超越所有现有开源模型
- 适合低资源语言研究者和多语言AI开发者
亚美尼亚语是一种形态复杂且资源匮乏的语言,其预训练数据稀缺,且尚未有开源亚美尼亚语大模型发布完整数据与训练方法。为此,我们构建并发布了两个数据集:ArmWeb为437万篇经过严格验证的亚美尼亚语新闻文档集合;ArmSTEM为包含37.3万道英亚双语数理题及其逐步解答的平行语料,翻译后通过答案保持型LLM判断与人工评估双重验证。在Gemma-4-E4B基础上继续预训练得到arm-gemma-e4b,该模型在性能上超过所有现有开源亚美尼亚语模型,并优于其原始基线模型,是首个提供完整训练数据与流程的开源亚美尼亚语大模型。消融实验表明,仅使用新闻数据微调可提升流畅性但削弱知识能力,这一现象也存在于现有模型中;而少量经验证的数理题数据可有效逆转知识损失。此外,我们发现最大公开亚美尼亚语语料库与网络来源的评测集存在高度重叠,甚至在FineWeb-2中出现训练集与测试集自重叠。所有数据、模型与代码均已公开。
原文摘要 · Abstract (English)
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。