arXiv:2410.16153cs.CLcs.CV2024-10被引 76

Pangea模型支持39种语言的多模态任务,打破英语主导局面。

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

  • 构建600万条跨语言指令数据集,融合人工与机器翻译。
  • 在47种语言上测试,性能显著优于现有开源多模态模型。
  • 数据、代码、模型全开源,适合追求语言公平的研究者。

尽管多模态大模型取得进展,但其发展主要聚焦于英语和西方数据集,导致全球多数语言和文化背景被忽视。本文提出Pangea,一个基于涵盖39种语言的600万条指令数据集PangeaIns训练的多语言多模态大模型。PangeaIns包含高质量英文指令、精心机器翻译的指令以及具有文化相关性的多模态任务,以确保跨文化覆盖。为全面评估模型能力,我们构建了PangeaBench,一个涵盖14个数据集、覆盖47种语言的综合评测套件。结果表明,Pangea在多语言和多元文化场景中显著优于现有开源模型。消融实验揭示英文数据比例、语言流行度及多模态训练样本数量对整体性能的关键影响。我们完全开源数据、代码和训练权重,推动更具包容性与鲁棒性的多语言多模态模型发展,促进更广泛的语言与文化平等与可及性。

原文摘要 · Abstract (English)

Despite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultural contexts underrepresented. This paper introduces Pangea, a multilingual multimodal LLM trained on PangeaIns, a diverse 6M instruction dataset spanning 39 languages. PangeaIns features: 1) high-quality English instructions, 2) carefully machine-translated instructions, and 3) culturally relevant multimodal tasks to ensure cross-cultural coverage. To rigorously assess models' capabilities, we introduce PangeaBench, a holistic evaluation suite encompassing 14 datasets covering 47 languages. Results show that Pangea significantly outperforms existing open-source models in multilingual settings and diverse cultural contexts. Ablation studies further reveal the importance of English data proportions, language popularity, and the number of multimodal training samples on overall performance. We fully open-source our data, code, and trained checkpoints, to facilitate the development of inclusive and robust multilingual MLLMs, promoting equity and accessibility across a broader linguistic and cultural spectrum.

多模态多语言开放数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。