首个中英双语多模态认知评估基准,诊断视觉语言模型真实推理能力。
Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

- 基于布卢姆认知分类,设计六层认知任务评估模型能力。
- 发现顶尖模型在事实记忆与创意生成上表现显著薄弱。
- 揭示中英文跨语言推理差距,推动更公平的多模态模型发展。
尽管视觉语言模型(VLMs)进展迅速,但缺乏能严谨诊断其真实推理能力并推动类人多模态智能发展的评估基准。现有评测多聚焦零散任务,掩盖了关键认知缺陷,难以指导针对性改进。为此,我们推出BloomBench,作为Almieyar系列的一部分,是首个基于人类认知、支持英语-阿拉伯语双语的多模态评估基准。它依据布卢姆分类法,通过精心设计的图像-问题-回答任务,系统评估六个认知层级(记忆、理解、应用、分析、评价、创造)。该基准采用半自动化流程构建,并通过分层混合质量保障协议验证,确保可扩展性、文化包容性与语言准确性。基于此框架,我们对主流VLMs进行了全面分析,发现其存在显著认知不对称:虽在语义理解上达到较高性能,但在事实记忆和创造性合成上表现明显不足。这表明当前通用多模态能力掩盖了特定认知层的深层局限。此外,研究揭示阿拉伯语与英语间存在显著性能差距,暴露了当前跨语言多模态推理的不足。这些发现为构建更符合认知规律且更具包容性的VLMs奠定了基础。基准框架与数据集已开源:https://github.com/qcri/Almieyar-Oryx-BloomBench。
原文摘要 · Abstract (English)
Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence. Most existing evaluations focus on piecemeal or disconnected tasks, obscuring critical cognitive weaknesses and providing little insight for targeted improvement. To address this gap, we introduce BloomBench, part of the Almieyar benchmarking series, the first cognitively human-grounded, bilingual (English-Arabic) multimodal benchmark for VLMs. Grounded in Bloom's Taxonomy, BloomBench systematically evaluates six levels of cognition (Remember, Understand, Apply, Analyze, Evaluate, Create) through carefully designed image-question-answer tasks. Built with a semi-automated pipeline and validated through a stratified hybrid quality assurance protocol, it ensures scalability, cultural inclusivity, and linguistic fidelity. Leveraging this framework, we conduct a comprehensive study of state-of-the-art VLMs to diagnose their cognitive profiles. Our analysis reveals a sharp cognitive asymmetry: while state-of-the-art models achieve strong performance ceilings in semantic understanding, they struggle substantially with factual recall and creative synthesis. This demonstrates that current general multimodal proficiency masks deeper limitations in specific cognitive layers. Furthermore, our study highlights a critical performance gap between Arabic and English, exposing limitations in current cross-lingual multimodal reasoning. These findings establish a foundation for developing more cognitively aligned and inclusive VLMs. The benchmark framework and dataset is available at: https://github.com/qcri/Almieyar-Oryx-BloomBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。