为阿富汗500万使用者的南方乌兹别克语构建首个机器翻译资源。
Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek
- 构建997句的FLORES+开发集与近4万句平行语料。
- 训练出专用模型lutfiy,提升南方乌兹别克语翻译性能。
- 开源全部数据与工具,助力低资源语言研究。
南方乌兹别克语(uzs)是阿富汗约500万人使用的突厥语变体,其语音、词汇和拼写与北方乌兹别克语(uzn)差异显著。尽管使用者众多,该语言在自然语言处理中仍严重缺乏资源。本文首次为南方乌兹别克语构建机器翻译资源,包括一个997句的FLORES+开发集、39,994句来自词典、文学和网络来源的平行语料,以及一个微调后的NLLB-200模型(lutfiy)。我们还提出一种后处理方法,用于恢复阿拉伯字母中的半空格字符,改善形态边界识别。所有数据集、模型和工具均公开发布,以支持南方乌兹别克语及其他低资源语言的研究。
原文摘要 · Abstract (English)
Southern Uzbek (uzs) is a Turkic language variety spoken by around 5 million people in Afghanistan and differs significantly from Northern Uzbek (uzn) in phonology, lexicon, and orthography. Despite the large number of speakers, Southern Uzbek is underrepresented in natural language processing. We present new resources for Southern Uzbek machine translation, including a 997-sentence FLORES+ dev set, 39,994 parallel sentences from dictionary, literary, and web sources, and a fine-tuned NLLB-200 model (lutfiy). We also propose a post-processing method for restoring Arabic-script half-space characters, which improves handling of morphological boundaries. All datasets, models, and tools are released publicly to support future work on Southern Uzbek and other low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。