arXiv:2411.16508cs.CVcs.CL2024-11CVPR被引 74

评测100种语言的多模态模型,推动文化多样性与低资源语言支持。

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

  • 构建包含100种语言的跨文化多模态评测基准ALM-bench。
  • 覆盖13类文化主题,测试模型在视觉与语言推理中的表现差异。
  • 面向关注全球包容性、文化敏感性的研究者与开发者。

现有大型多模态模型(LMMs)通常仅聚焦少数地区和语言。随着模型性能提升,确保其理解文化背景、尊重地方敏感性并支持低资源语言变得愈发重要,同时需有效整合对应视觉线索。为此,我们提出「所有语言都重要」基准(ALM-bench),这是迄今最大且最全面的跨100种语言的LMM评估体系。ALM-bench通过包含多种语言的图文对,挑战模型在文化多样性场景下的理解与推理能力,涵盖大量传统上被忽视的低资源语言。该基准采用多种题型(判断题、选择题、开放问答),分为短答与长答两类,全面评估模型在不同难度下视觉与语言推理的表现。内容基于13个文化维度精心设计,包括传统习俗、仪式、名人、节庆等。ALM-bench不仅为前沿开源与闭源模型提供严格评测平台,更强调文化与语言包容性的重要性,推动开发能服务全球多元人群的模型。基准已公开可用。

原文摘要 · Abstract (English)

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model's ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available.

多模态语言多样性文化评测低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。