用阿拉伯语数据增强马耳他语NLP,提升小语种处理效果
Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data
- 通过音译和机器翻译将阿拉伯语文本映射到马耳他语
- 实验证明阿拉伯语增强可显著提升马耳他语模型性能
- 适合关注小语种、跨语言迁移的研究者
马耳他语是一种独特的闪米特语,受罗曼语和日耳曼语(尤其是意大利语和英语)强烈影响,尽管其词源为闪米特语,但书写系统采用拉丁字母,与最近的亲属语言阿拉伯语存在明显差异。本文探讨能否利用阿拉伯语资源,通过跨语言数据增强技术支持马耳他语自然语言处理。研究比较了多种阿拉伯语文本与马耳他语对齐策略,包括不同音译方案和机器翻译方法,并提出新型音译系统以更准确反映马耳他语拼写。评估了这些增强方法在单语和多语模型上的效果,结果表明基于阿拉伯语的数据增强能显著改善马耳他语任务表现。
原文摘要 · Abstract (English)
Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin script, creating a gap between it and its closest linguistic relatives in Arabic. In this paper, we explore whether Arabic-language resources can support Maltese natural language processing (NLP) through cross-lingual augmentation techniques. We investigate multiple strategies for aligning Arabic textual data with Maltese, including various transliteration schemes and machine translation (MT) approaches. As part of this, we also introduce novel transliteration systems that better represent Maltese orthography. We evaluate the impact of these augmentations on monolingual and mutlilingual models and demonstrate that Arabic-based augmentation can significantly benefit Maltese NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。