arXiv:2505.08910cs.CVcs.CL2025-05中稿 · CVPR被引 3

构建多语言视觉语言模型Maya,提升低资源语言理解能力

Behind Maya: Building a Multilingual Vision Language Model

  • 基于LLaVA数据集扩展八种语言的图文预训练数据
  • 模型支持八种语言,在跨文化场景中表现更优
  • 开源模型与数据,适合多语言研究者使用

近年来,大型视觉语言模型(VLM)发展迅速,在主流语言的学术基准上表现优异,但对低资源语言和多样化文化背景的适应性不足。为解决这一问题,我们提出Maya——一个开源的多语言视觉语言模型。主要贡献包括:1)基于LLaVA预训练数据集构建的八种语言图文预训练数据集;2)支持这八种语言的多语言图像-文本模型,增强了在视觉语言任务中的文化与语言理解能力。代码已公开于https://github.com/nahidalam/maya。

原文摘要 · Abstract (English)

In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily in widely spoken languages but lack performance on low-resource languages and varied cultural contexts. To address these limitations, we introduce Maya, an open-source Multilingual VLM. Our contributions are: 1) a multilingual image-text pretraining dataset in eight languages, based on the LLaVA pretraining dataset; and 2) a multilingual image-text model supporting these languages, enhancing cultural and linguistic comprehension in vision-language tasks. Code available at https://github.com/nahidalam/maya.

多语言模型视觉语言开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。