arXiv:2505.17021cs.CV2025-05被引 6

首个评估阿拉伯语多模态推理的基准,覆盖11个领域

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

  • 构建11个领域的阿拉伯语多模态推理数据集
  • 包含1356组图文样本与5119条人工标注推理步骤
  • 适合关注低资源语言和文化感知AI的研究者

随着大型多模态模型(LMMs)能力提升,对其推理过程的评估日益重要。然而,现有基准大多聚焦英语,忽视了阿拉伯语等具有丰富语言与文化背景的语言。为此,我们提出首个面向阿拉伯语多模态推理的综合性基准ARB,涵盖视觉推理、文档理解、OCR、科学分析及文化解读等11个领域。ARB包含1,356组多模态样本与5,119条人工标注的推理步骤及对应操作。我们评估了12个主流开源与闭源LMMs,发现其在连贯性、忠实性和文化契合度方面仍存在显著挑战。ARB为诊断低资源语言中的多模态推理提供了结构化框架,推动更包容、透明且具文化意识的AI发展。我们已公开发布基准数据、评分标准与评估工具套件,代码见:https://github.com/mbzuai-oryx/ARB。

原文摘要 · Abstract (English)

As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks remain focused on English, overlooking languages with rich linguistic and cultural contexts, such as Arabic. To address this gap, we introduce the Comprehensive Arabic Multimodal Reasoning Benchmark (ARB), the first benchmark designed to evaluate step-by-step reasoning in Arabic across both textual and visual modalities. ARB spans 11 diverse domains, including visual reasoning, document understanding, OCR, scientific analysis, and cultural interpretation. It comprises 1,356 multimodal samples paired with 5,119 human-curated reasoning steps and corresponding actions. We evaluated 12 state-of-the-art open- and closed-source LMMs and found persistent challenges in coherence, faithfulness, and cultural grounding. ARB offers a structured framework for diagnosing multimodal reasoning in underrepresented languages and marks a critical step toward inclusive, transparent, and culturally aware AI systems. We release the benchmark, rubric, and evaluation suit to support future research and reproducibility. Code available at: https://github.com/mbzuai-oryx/ARB

多模态推理阿拉伯语低资源语言文化感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。