研究发现语言类型特征影响多语言翻译质量,尤其在大模型时代仍具关键作用。
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
- 分析NLLB-200和Tower+两大主流多语言模型,验证语言类型对翻译质量的影响。
- 即使控制资源量与书写系统后,目标语言的类型学特征仍显著影响翻译效果。
- 某些语言更适合采用非标准解码策略,可提升翻译表现,适合模型优化研究者参考。
尽管多语言建模取得显著进展,但不同语言间的翻译质量差异依然明显。除了训练资源不均外,语言的类型学特性也被认为是决定建模难度的内在因素。现有研究多基于小型单语或双语模型,而本文扩展到两个大型预训练多语言翻译模型——代表编码器-解码器架构的NLLB-200和代表解码器仅架构的Tower+。在涵盖广泛语言的基础上,研究发现目标语言的类型学特征对两类模型的翻译质量均有显著影响,即便在控制资源丰富度和书写系统等外部因素后依然成立。此外,具有特定类型学属性的语言在更宽泛的输出空间搜索中受益更大,表明其可能从非标准左向右束搜索的解码策略中获益。为促进该领域研究,本文发布了针对FLORES+评测基准中212种语言的细粒度类型学属性数据集。
原文摘要 · Abstract (English)
Despite major advances in multilingual modeling, large quality disparities persist across languages. Besides the obvious impact of uneven training resources, typological properties have also been proposed to determine the intrinsic difficulty of modeling a language. The existing evidence, however, is mostly based on small monolingual language models or bilingual translation models trained from scratch. We expand on this line of work by analyzing two large pre-trained multilingual translation models, NLLB-200 and Tower+, which are state-of-the-art representatives of encoder-decoder and decoder-only machine translation, respectively. Based on a broad set of languages, we find that target language typology drives translation quality of both models, even after controlling for more trivial factors, such as data resourcedness and writing script. Additionally, languages with certain typological properties benefit more from a wider search of the output space, suggesting that such languages could profit from alternative decoding strategies beyond the standard left-to-right beam search. To facilitate further research in this area, we release a set of fine-grained typological properties for 212 languages of the FLORES+ MT evaluation benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。