研究发现,语言差异比数据量和模型大小更影响低资源语言翻译效果。
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
- 对比两种巴西原住民语言的微调翻译性能,分析训练因素影响。
- 不同语言结构差异导致翻译效果显著不同,其他因素影响有限。
- 对濒危语言翻译有指导意义,适合语言保护与计算语言学研究者。
使用少量数据微调预训练语言模型是为超低资源语言(如濒危原住民语言)创建翻译系统的一种常用方法。然而,已有研究显示,采用相似方法和数据生成的翻译器性能却存在显著差异。本文系统探究了性能差异的可能原因,包括数据清洗方式、预训练模型局限性、基础模型规模及训练数据量,覆盖双向翻译任务。研究以两种结构特征差异显著的巴西原住民语言为对象,结果表明上述训练因素影响极小或无明显影响,暗示语言间的本质差异可能在微调翻译能力中起关键作用。
原文摘要 · Abstract (English)
Finetuning pre-trained language models with small amounts of data is a commonly-used method to create translators for ultra-low resource languages such as endangered Indigenous languages. However, previous works have reported substantially different performances with translators created using similar methodology and data. In this work we systematically explored possible causes of the performance difference, aiming to determine whether it was a product of different cleaning procedures, limitations of the pre-trained models, the size of the base model, or the size of the training dataset, studying both directions of translation. Our studies, using two Brazilian Indigenous languages, related but with significant structural linguistic characteristics, indicated none or very limited influence from those training factors, suggesting differences between languages may play a significant role in the ability to produce translators by fine-tuning pre-trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。