首个非洲多语言文化问答基准,覆盖15种语言7500个问答对。
Afri-MCQA: Multimodal Cultural Question Answering for African Languages
- 由母语者构建跨文本与语音的多模态问答数据集。
- 开源模型在原生语言问答中准确率接近零,语音识别更差。
- 适合关注非洲语言AI、跨文化迁移与语音优先模型的研究者。
非洲拥有全球超过三分之一的语言,但在人工智能研究中仍被严重忽视。我们提出Afri-MCQA,首个覆盖15种非洲语言(来自12个国家)的多语言文化问答基准,包含7500个问答对。该基准提供平行的英文-非洲语言文本与语音问答对,全部由母语者创建。在Afri-MCQA上评估大型语言模型发现,开放权重模型在所评估文化中表现不佳,以原生语言或语音提问时,开放式视觉问答准确率几乎为零。为评估语言能力,我们设计了控制实验以分离语言与文化知识的影响,结果显示原生语言与英语之间在文本和语音任务中均存在显著性能差距。这些结果凸显了语音优先方法、基于文化背景的预训练以及跨语言文化迁移的必要性。为支持非洲语言的包容性多模态AI发展,我们已将Afri-MCQA以学术许可或CC BY-NC 4.0协议发布于HuggingFace(https://huggingface.co/datasets/Atnafu/Afri-MCQA)。
原文摘要 · Abstract (English)
Africa is home to over one-third of the world's languages, yet remains underrepresented in AI research. We introduce Afri-MCQA, the first Multilingual Cultural Question-Answering benchmark covering 7.5k Q&A pairs across 15 African languages from 12 countries. The benchmark offers parallel English-African language Q&A pairs across text and speech modalities and was entirely created by native speakers. Benchmarking large language models (LLMs) on Afri-MCQA shows that open-weight models perform poorly across evaluated cultures, with near-zero accuracy on open-ended VQA when queried in native language or speech. To evaluate linguistic competence, we include control experiments meant to assess this specific aspect separate from cultural knowledge, and we observe significant performance gaps between native languages and English for both text and speech. These findings underscore the need for speech-first approaches, culturally grounded pretraining, and cross-lingual cultural transfer. To support more inclusive multimodal AI development in African languages, we release our Afri-MCQA under academic license or CC BY-NC 4.0 on HuggingFace (https://huggingface.co/datasets/Atnafu/Afri-MCQA)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。