首个兼顾英阿双语的多模态大模型,提升阿拉伯语视觉理解能力。
AIN: The Arabic INclusive Large Multimodal Model
- 构建360万条高质量英阿多模态数据,训练双语多模态模型。
- 7B版本在38个子领域平均领先GPT-4o 3.4%,跨域表现优异。
- 适合需阿拉伯语多模态应用的开发者与研究者使用。
随着大语言模型(LLMs)向多模态大模型(LMMs)演进,英语和中文等高资源语言已取得显著进展。尽管阿拉伯语大模型发展迅速,但阿拉伯语多模态模型仍处于探索阶段,通常局限于语言和视觉理解的少数方面。为此,我们提出AIN——阿拉伯语包容性多模态模型,旨在覆盖多样应用场景。AIN是英阿双语多模态模型,利用精心构建的360万条高质量阿拉伯语-英语多模态数据样本进行训练。AIN在阿拉伯语任务上达到顶尖水平,同时具备强大的英语视觉理解能力。在包含多图像理解、复杂视觉感知、手写文档识别、视频理解、医学影像、植物病害及遥感土地利用理解等38个子领域的最新CAMEL-Bench基准测试中,7B版本模型在8个领域平均超越GPT-4o 3.4%。AIN的卓越性能标志着为阿拉伯语使用者提供先进多模态生成式AI工具的重要进展。
原文摘要 · Abstract (English)
Amid the swift progress of large language models (LLMs) and their evolution into large multimodal models (LMMs), significant strides have been made in high-resource languages such as English and Chinese. While Arabic LLMs have seen notable progress, Arabic LMMs remain largely unexplored, often narrowly focusing on a few specific aspects of the language and visual understanding. To bridge this gap, we introduce AIN-the Arabic Inclusive Multimodal Model-designed to excel across diverse domains. AIN is an English-Arabic bilingual LMM designed to excel in English and Arabic, leveraging carefully constructed 3.6 million high-quality Arabic-English multimodal data samples. AIN demonstrates state-of-the-art Arabic performance, while also possessing strong English-language visual capabilities. On the recent CAMEL-Bench benchmark comprising 38 sub-domains including, multi-image understanding, complex visual perception, handwritten document understanding, video understanding, medical imaging, plant diseases, and remote sensing-based land use understanding, our AIN demonstrates strong performance with the 7B model outperforming GPT-4o by an absolute gain of 3.4% averaged over eight domains and 38 sub-domains. AIN's superior capabilities position it as a significant step toward empowering Arabic speakers with advanced multimodal generative AI tools across diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。