让英语训练的视觉语言模型学会多语言嵌入,提升跨模态理解能力。
xVLM2Vec: Adapting LVLM-based embedding models to multilinguality using Self-Knowledge Distillation
- 用自知识蒸馏方法将英语模型迁移到多语言任务。
- 在多语言图文数据上显著提升嵌入效果,尤其在低资源语言中。
- 提出首个多语言多模态嵌入评估基准,适合跨语言研究者使用。
当前大多数嵌入模型基于仅编码器的Transformer架构,以提取文本、图像等输入的稠密且有意义的表示。随着大语言模型的发展,利用这些大规模预训练模型提取嵌入成为可能。然而,现有研究主要聚焦于英语文本嵌入,而这些模型也主要在英语数据上训练。此外,支持多模态和多语言输入的模型极为稀少。为此,本文提出一种针对英语训练的大型视觉-语言模型(LVLM)的适配方法,以增强其在多语言和多模态嵌入上的表现。最后,我们设计并引入一个基准测试集,用于评估多语言与多模态嵌入模型的有效性。
原文摘要 · Abstract (English)
In the current literature, most embedding models are based on the encoder-only transformer architecture to extract a dense and meaningful representation of the given input, which can be a text, an image, and more. With the recent advances in language modeling thanks to the introduction of Large Language Models, the possibility of extracting embeddings from these large and extensively trained models has been explored. However, current studies focus on textual embeddings in English, which is also the main language on which these models have been trained. Furthermore, there are very few models that consider multimodal and multilingual input. In light of this, we propose an adaptation methodology for Large Vision-Language Models trained on English language data to improve their performance in extracting multilingual and multimodal embeddings. Finally, we design and introduce a benchmark to evaluate the effectiveness of multilingual and multimodal embedding models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。