评测阿拉伯语多模态模型在文化语境下的表现,涵盖图文问答与生成准确性。
ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

- 构建双任务评测框架:口语视觉问答与图像幻觉检测,及文本到图像的文化准确性评估。
- 14支团队参与,系统采用零样本提示、微调视觉语言模型等多样化方法。
- 首次系统性推动阿拉伯语多模态文化理解评测,适合关注AI文化偏见的研究者。
本文介绍 ImageEval 2026 共同任务中关于文化语境下阿拉伯语多模态评估的概述。该任务包含两个子任务:(i) AynVQA,覆盖英语与现代标准阿拉伯语(MSA)的口语视觉问答和基于图像的幻觉检测;(ii) CRAI-Bench,评估文本到图像生成的文化准确性。共有14支团队参与测试阶段,其中12支提交了系统描述论文。参赛系统采用多种方法,包括零样本提示、视觉语言模型微调、语音识别流水线、集成与分数校准。本文详细描述任务设置、数据集、评估流程及参与系统,并总结各赛道的主要结果。所有数据集与评估脚本已向研究社区开源。该共同任务凸显了阿拉伯语语音与图文推理中文化语境评估的挑战。
原文摘要 · Abstract (English)
We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic (MSA), and (ii) CRAI-Bench, evaluating the cultural accuracy of text-to-image generation. A total of 14 teams participated in the test phase, with 12 teams submitting system description papers. Participating systems used a range of approaches, including zero-shot prompting, fine-tuning of vision-language models, speech-recognition pipelines, ensembling, and score calibration. We describe the task setup, datasets, evaluation procedure, and participating systems, and summarize the main results across the different tracks. All datasets and evaluation scripts from the shared task are released to the research community. The shared task highlights the challenges of culturally grounded multimodal evaluation, particularly for Arabic speech and image-text reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。