arXiv:2510.16198cs.CL2025-10

构建埃及文化多模态数据集,填补中东非洲地区AI训练空白

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

  • 设计新采集流程,收集3000+张涵盖313个文化概念的图像
  • 在埃塞俄比亚文化数据集上CLIP模型零样本分类准确率仅21.2%Top-1
  • 适合研究文化偏见、视觉语言模型本地化及中东非洲文化研究者

尽管人工智能取得进展,但多模态文化多样性数据集仍十分有限,尤其在中东和非洲地区。本文提出EgMM-Corpus,一个专注于埃及文化的多模态数据集。通过设计并运行新的数据采集流程,我们收集了超过3,000张图像,覆盖313个概念,包括地标、食物和民俗。每个条目均经过人工验证以确保文化真实性和多模态一致性。EgMM-Corpus旨在为视觉-语言模型在埃及文化背景下的评估与训练提供可靠资源。我们进一步评估了对比语言-图像预训练(CLIP)在该数据集上的零样本性能,其分类任务中达到21.2%的Top-1准确率和36.4%的Top-5准确率。这些结果凸显了大规模视觉-语言模型中存在的文化偏见,并证明了EgMM-Corpus作为开发文化感知模型基准的重要性。

原文摘要 · Abstract (English)

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to Egyptian culture. By designing and running a new data collection pipeline, we collected over 3,000 images, covering 313 concepts across landmarks, food, and folklore. Each entry in the dataset is manually validated for cultural authenticity and multimodal coherence. EgMM-Corpus aims to provide a reliable resource for evaluating and training vision-language models in an Egyptian cultural context. We further evaluate the zero-shot performance of Contrastive Language-Image Pre-training CLIP on EgMM-Corpus, on which it achieves 21.2% Top-1 accuracy and 36.4% Top-5 accuracy in classification. These results underscore the existing cultural bias in large-scale vision-language models and demonstrate the importance of EgMM-Corpus as a benchmark for developing culturally aware models.

多模态数据集埃及文化视觉语言模型文化偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。