首个面向阿拉伯语的多模态模型评测基准,覆盖38个子领域。
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
- 构建涵盖8大领域的阿拉伯语多模态评测集
- 29,036道题经母语者人工校验,准确率仅62%(GPT-4o)
- 适合关注中东语言、跨语言通用性的研究者
近年来,大型多模态模型(LMMs)在视觉推理与理解任务上取得显著进展,催生了多种评估基准。然而,现有评测基准大多以英语为中心。本文提出CAMEL-Bench,首个针对阿拉伯语的综合性多模态模型评测基准,服务于超过4亿阿拉伯语使用者。该基准包含8个主要领域和38个子领域,涵盖多图像理解、复杂视觉感知、手写文档识别、视频理解、医学影像、植物病害及遥感土地利用分析等,全面评估模型在真实场景下的泛化能力。共包含约29,036道经过筛选并由母语者人工验证的问题,确保评估可靠性。我们对闭源(如GPT-4系列)和开源模型进行了评估,结果显示当前最佳模型仍需大幅提升,即便最先进的闭源模型GPT-4o整体得分也仅为62%。相关数据与评估脚本已开源。
原文摘要 · Abstract (English)
Recent years have witnessed a significant interest in developing large multimodal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led to the introduction of multiple LMM benchmarks to evaluate LMMs on different tasks. However, most existing LMM evaluation benchmarks are predominantly English-centric. In this work, we develop a comprehensive LMM evaluation benchmark for the Arabic language to represent a large population of over 400 million speakers. The proposed benchmark, named CAMEL-Bench, comprises eight diverse domains and 38 sub-domains including, multi-image understanding, complex visual perception, handwritten document understanding, video understanding, medical imaging, plant diseases, and remote sensing-based land use understanding to evaluate broad scenario generalizability. Our CAMEL-Bench comprises around 29,036 questions that are filtered from a larger pool of samples, where the quality is manually verified by native speakers to ensure reliable model assessment. We conduct evaluations of both closed-source, including GPT-4 series, and open-source LMMs. Our analysis reveals the need for substantial improvement, especially among the best open-source models, with even the closed-source GPT-4o achieving an overall score of 62%. Our benchmark and evaluation scripts are open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。