首个专为意第绪语设计的80亿参数语言模型,填补低资源语言NLP空白。
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
- 基于高质量语料Oytser和多任务基准Kashes,微调Llama 3.1 8B构建
- 在多项任务上超越同规模开源基线,显著提升意第绪语理解能力
- 为历史丰富但数字匮乏的语言提供可复用的建模范式
我们提出MameLoshnLM,首个针对意第绪语的开源80亿参数语言模型。尽管意第绪语拥有丰富的文本传统,其有限的数字存在及可靠评估资源的缺乏制约了该语言的自然语言处理进展。现有多语言语料库和基准常作为不良代理,包含大量噪声、机器翻译和误分类文本。为此,我们引入Oytser——一个结合当代网络来源与文学材料的高质量意第绪语预训练语料库,以及涵盖翻译、语言分析、信息抽取与语言理解的多任务基准Kashes。利用这些资源,我们对Llama 3.1 8B进行继续预训练,得到MameLoshnLM。在基准各项任务中,该模型优于同等规模的开源基线。分析表明,这种提升不仅是量化的:相较于通用多语言模型,MameLoshnLM更准确捕捉语言特有的词汇与形态模式,揭示出噪声网络规模多语言数据在低资源语言上的系统性缺陷。结果为意第绪语自然语言处理奠定基础,并为历史上丰富但数字代表性不足的语言提供实际建模模板。
原文摘要 · Abstract (English)
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。