首个评测波斯语视觉语言模型的多模态教育数据集
MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- 构建涵盖1-12年级的波斯语多模态试题库
- 含7500道波斯语题,支持跨语言性能对比
- 专为评估模型理解力与幻觉倾向设计
大型视觉语言模型(VLMs)近期进展主要集中在英语,对其他语言关注不足。为此,我们推出MEENA(又称PersianMMMU),首个用于评估波斯语VLMs在科学、推理和人类理解任务上的基准数据集。该数据集包含约7,500道波斯语题目和3,000道英语题目,覆盖数学、物理、图表、图表、波斯艺术与文学等广泛主题。关键特征包括:(1) 跨小学至高中多个教育阶段的多样化学科内容;(2) 丰富元数据,含难度等级与描述性答案;(3) 原创波斯语内容,保留文化细节;(4) 双语结构以评估跨语言表现;(5) 多项实验验证模型整体性能、图像关注度及幻觉生成倾向。我们期望该基准推动VLM能力超越英语范畴。
原文摘要 · Abstract (English)
Recent advancements in large vision-language models (VLMs) have primarily focused on English, with limited attention given to other languages. To address this gap, we introduce MEENA (also known as PersianMMMU), the first dataset designed to evaluate Persian VLMs across scientific, reasoning, and human-level understanding tasks. Our dataset comprises approximately 7,500 Persian and 3,000 English questions, covering a wide range of topics such as reasoning, mathematics, physics, diagrams, charts, and Persian art and literature. Key features of MEENA include: (1) diverse subject coverage spanning various educational levels, from primary to upper secondary school, (2) rich metadata, including difficulty levels and descriptive answers, (3) original Persian data that preserves cultural nuances, (4) a bilingual structure to assess cross-linguistic performance, and (5) a series of diverse experiments assessing various capabilities, including overall performance, the model's ability to attend to images, and its tendency to generate hallucinations. We hope this benchmark contributes to enhancing VLM capabilities beyond English.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。