用图文对训练让多语言文本表示自动对齐,无需双语数据。
Cross-Lingual Representation Alignment Through Contrastive Image-Caption Tuning
- 利用图文数据集进行对比学习,实现跨语言文本表示对齐。
- 无需双语数据,新语言可事后加入对齐,支持低资源语言。
- 对齐后的表示可用于跨语言理解与双语检索任务。
多语言句子表示的对齐通常依赖双语语料来弥合语言间的差距。本文探讨是否可用视觉信息替代双语数据来实现这一目标。图像描述数据集无需多语言专业知识即可轻松构建,为低资源语言提供了更高效的解决方案。研究发现,通过多语言图文对齐可隐式实现不同语言间文本表示的对齐;在预训练中未见过的语言也可在事后加入对齐过程;这些对齐后的表示可直接用于跨语言自然语言理解(NLU)和双语语料检索任务。
原文摘要 · Abstract (English)
Multilingual alignment of sentence representations has mostly required bitexts to bridge the gap between languages. We investigate whether visual information can bridge this gap instead. Image caption datasets are very easy to create without requiring multilingual expertise, so this offers a more efficient alternative for low-resource languages. We find that multilingual image-caption alignment can implicitly align the text representations between languages, languages unseen by the encoder in pretraining can be incorporated into this alignment post-hoc, and these aligned representations are usable for cross-lingual Natural Language Understanding (NLU) and bitext retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。