用少量卢森堡语+等量德法语训练模型,提升低资源语言生成效果
Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy
- 混合卢森堡语、德语和法语数据,均衡训练多语言模型
- 新构建的LuxGen基准测试显示模型在生成质量上优于单语和大型多语模型
- 适合低资源语言研究者及多语言自然语言处理开发者参考
本文针对代表性不足语言的模型开发挑战,聚焦卢森堡语。尽管卢森堡语持续发展,但其数字数据稀缺,且受卢森堡多语言环境加剧。我们提出基于T5架构的新文本生成模型,将有限卢森堡语数据与同等规模和类型(种类)的德语和法语数据结合。假设该模型在跨语言迁移学习方面表现更优,能超越单语和大型多语模型。为验证此假设,本研究探讨了多语言与单语言训练对卢森堡语生成任务的相对优势。评估中引入首个卢森堡语文本生成基准——LuxGen。
原文摘要 · Abstract (English)
This paper addresses the challenges in developing language models for less-represented languages, with a focus on Luxembourgish. Despite its active development, Luxembourgish faces a digital data scarcity, exacerbated by Luxembourg's multilingual context. We propose a novel text generation model based on the T5 architecture, combining limited Luxembourgish data with equal amounts, in terms of size and type, of German and French data. We hypothesise that a model trained on Luxembourgish, German, and French will improve the model's cross-lingual transfer learning capabilities and outperform monolingual and large multilingual models. To verify this, the study at hand explores whether multilingual or monolingual training is more beneficial for Luxembourgish language generation. For the evaluation, we introduce LuxGen, a text generation benchmark that is the first of its kind for Luxembourgish.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。