arXiv:2502.00094cs.CVcs.AI2025-02被引 13

首个兼顾英阿双语的多模态大模型,提升阿拉伯语视觉理解能力。

AIN: The Arabic INclusive Large Multimodal Model

  • 构建360万条高质量英阿多模态数据,训练双语多模态模型。
  • 7B版本在38个子领域平均领先GPT-4o 3.4%,跨域表现优异。
  • 适合需阿拉伯语多模态应用的开发者与研究者使用。

随着大语言模型(LLMs)向多模态大模型(LMMs)演进,英语和中文等高资源语言已取得显著进展。尽管阿拉伯语大模型发展迅速,但阿拉伯语多模态模型仍处于探索阶段,通常局限于语言和视觉理解的少数方面。为此,我们提出AIN——阿拉伯语包容性多模态模型,旨在覆盖多样应用场景。AIN是英阿双语多模态模型,利用精心构建的360万条高质量阿拉伯语-英语多模态数据样本进行训练。AIN在阿拉伯语任务上达到顶尖水平,同时具备强大的英语视觉理解能力。在包含多图像理解、复杂视觉感知、手写文档识别、视频理解、医学影像、植物病害及遥感土地利用理解等38个子领域的最新CAMEL-Bench基准测试中,7B版本模型在8个领域平均超越GPT-4o 3.4%。AIN的卓越性能标志着为阿拉伯语使用者提供先进多模态生成式AI工具的重要进展。

原文摘要 · Abstract (English)

Amid the swift progress of large language models (LLMs) and their evolution into large multimodal models (LMMs), significant strides have been made in high-resource languages such as English and Chinese. While Arabic LLMs have seen notable progress, Arabic LMMs remain largely unexplored, often narrowly focusing on a few specific aspects of the language and visual understanding. To bridge this gap, we introduce AIN-the Arabic Inclusive Multimodal Model-designed to excel across diverse domains. AIN is an English-Arabic bilingual LMM designed to excel in English and Arabic, leveraging carefully constructed 3.6 million high-quality Arabic-English multimodal data samples. AIN demonstrates state-of-the-art Arabic performance, while also possessing strong English-language visual capabilities. On the recent CAMEL-Bench benchmark comprising 38 sub-domains including, multi-image understanding, complex visual perception, handwritten document understanding, video understanding, medical imaging, plant diseases, and remote sensing-based land use understanding, our AIN demonstrates strong performance with the 7B model outperforming GPT-4o by an absolute gain of 3.4% averaged over eight domains and 38 sub-domains. AIN's superior capabilities position it as a significant step toward empowering Arabic speakers with advanced multimodal generative AI tools across diverse applications.

多模态阿拉伯语双语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。