arXiv:2411.14647cs.CL2024-11被引 3

构建首个乌克兰语多模态评估基准,推动低资源语言模型发展

Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains

  • 基于乌克兰高考题构建4300+多模态题目库,覆盖12个学科
  • 仅少数模型性能超基线,暴露乌克兰语多模态能力短板
  • 首次评测乌克兰语图文生成与文化知识理解,适合低资源语言研究者

尽管英语多模态模型的评估已较为成熟,但低资源和中资源语言仍缺乏系统性评测。本文提出ZNO-Vision,一个源自标准化大学入学考试(ZNO)的乌克兰语多模态基准,包含超过4300道专家设计的题目,覆盖数学、物理、化学及人文学科等12个领域。我们评估了开源模型与API服务的表现,发现仅有少数模型优于基线。同时,首次对乌克兰语多模态文本生成进行评估:在Multi30K-UK数据集上测量图像描述质量,将VQA基准翻译为乌克兰语并量化性能下降;还从文化角度测试了模型对乌克兰民族菜肴的知识掌握情况。本工作有望推动乌克兰语多模态生成能力发展,其方法亦可推广至其他低资源语言。

原文摘要 · Abstract (English)

While the evaluation of multimodal English-centric models is an active area of research with numerous benchmarks, there is a profound lack of benchmarks or evaluation suites for low- and mid-resource languages. We introduce ZNO-Vision, a comprehensive multimodal Ukrainian-centric benchmark derived from standardized university entrance examination (ZNO). The benchmark consists of over 4,300 expert-crafted questions spanning 12 academic disciplines, including mathematics, physics, chemistry, and humanities. We evaluated the performance of both open-source models and API providers, finding that only a handful of models performed above baseline. Alongside the new benchmark, we performed the first evaluation study of multimodal text generation for the Ukrainian language: we measured caption generation quality on the Multi30K-UK dataset, translated the VQA benchmark into Ukrainian, and measured performance degradation relative to original English versions. Lastly, we tested a few models from a cultural perspective on knowledge of national cuisine. We believe our work will advance multimodal generation capabilities for the Ukrainian language and our approach could be useful for other low-resource languages.

多模态低资源语言评估基准乌克兰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。