研究跨语言迁移中并行数据的作用,发现机器翻译数据平均最优,但真实语料在部分语言上更优。
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
- 用并行数据迁移已训练编码器,探索语言和领域影响。
- 机器翻译任务数据平均表现最佳,真实语料在部分语言上超越它。
- 多数语言从多语言训练中获益,适合多语言视觉语言任务研究者。
大多数预训练的视觉-语言(VL)模型及下游任务数据仅提供英文。因此,多语言VL任务通常通过跨语言迁移解决:微调多语言预训练模型或使用并行数据转移文本编码器。本文研究另一种方法:利用并行数据迁移已训练的编码器。我们考察了并行数据的领域和语言数量的影响,这些因素在先前工作中未受关注。结果表明,即使机器翻译的任务数据在平均表现上最优,某些语言的真实平行数据(如字幕)仍表现更佳。此外,我们证明大多数语言都能从多语言训练中受益。
原文摘要 · Abstract (English)
Most pre-trained Vision-Language (VL) models and training data for the downstream tasks are only available in English. Therefore, multilingual VL tasks are solved using cross-lingual transfer: fine-tune a multilingual pre-trained model or transfer the text encoder using parallel data. We study the alternative approach: transferring an already trained encoder using parallel data. We investigate the effect of parallel data: domain and the number of languages, which were out of focus in previous work. Our results show that even machine-translated task data are the best on average, caption-like authentic parallel data outperformed it in some languages. Further, we show that most languages benefit from multilingual training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。