arXiv:2603.29467cs.CLcs.AI2026-03被引 1

M-MiniGPT4通过翻译数据实现11种语言的视觉语言理解,性能超越同类模型。

M-MiniGPT4: Multilingual VLLM Alignment via Translated Data

  • 融合原生多语言与翻译数据,提升跨语言对齐能力。
  • 在11种语言的MMMU基准上达36%准确率,领先同规模模型。
  • 适合低资源语言研究,开源数据与代码推动多语种发展。

本文提出多语言视觉语言模型M-MiniGPT4,具备11种语言的强视觉语言理解能力。通过混合使用原生多语言数据与翻译数据,优化了MiniGPT4架构的多语言表现。此外,我们设计了基于平行语料库的多语言对齐训练阶段,进一步增强模型跨语言能力。M-MiniGPT4在多语言MMMU基准上取得36%的准确率,优于同参数量级的现有模型,包括本工作完成后发布的部分基础模型。我们开源了模型、代码及翻译数据集,以支持低资源和多语言场景下的后续研究。

原文摘要 · Abstract (English)

This paper presents a Multilingual Vision Large Language Model, named M-MiniGPT4. Our model exhibits strong vision-language understanding (VLU) capabilities across 11 languages. We utilize a mixture of native multilingual and translated data to push the multilingual VLU performance of the MiniGPT4 architecture. In addition, we propose a multilingual alignment training stage that uses parallel text corpora to further enhance the multilingual capabilities of our model. M-MiniGPT4 achieves 36% accuracy on the multilingual MMMU benchmark, outperforming state-of-the-art models in the same weight class, including foundation models released after the majority of this work was completed. We open-source our models, code, and translated datasets to facilitate future research in low-resource and multilingual settings.

多语言视觉语言对齐训练开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。