首个聚焦印度次大陆的多语言视觉语言模型评测基准
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
- 构建覆盖10种印度语言的多模态评测集
- 发现主流模型在文化多样性场景下表现显著下降
- 适合关注公平性与跨文化AI的研究者使用
视觉语言模型在多模态任务中展现出强大泛化能力,但现有评测基准大多以西方为中心,难以评估其在文化多样性和多语言环境下的表现。为此,我们提出IndicVisionBench,首个聚焦印度次大陆的大规模评测基准。涵盖英语及10种印度语言,覆盖光学字符识别(OCR)、多模态机器翻译(MMT)和视觉问答(VQA)3类任务,包含6种问题类型。最终基准包含约5000张图像和37,000+个问答对,覆盖13个文化相关主题。同时发布10种印地语系语言的配对平行标注语料库,可用于分析视觉语言模型中的文化和语言偏见。我们评估了8种模型,包括闭源系统和开源中大型模型,实验显示存在显著性能差距,凸显当前模型在文化多样性场景下的局限性。通过聚焦文化多样性和多语言性,IndicVisionBench建立了一个可复现的评估框架,推动更包容的多模态研究。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diverse and multilingual settings. To address this gap, we introduce IndicVisionBench, the first large-scale benchmark centered on the Indian subcontinent. Covering English and 10 Indian languages, our benchmark spans 3 multimodal tasks, including Optical Character Recognition (OCR), Multimodal Machine Translation (MMT), and Visual Question Answering (VQA), covering 6 kinds of question types. Our final benchmark consists of a total of ~5K images and 37K+ QA pairs across 13 culturally grounded topics. In addition, we release a paired parallel corpus of annotations across 10 Indic languages, creating a unique resource for analyzing cultural and linguistic biases in VLMs. We evaluate a broad spectrum of 8 models, from proprietary closed-source systems to open-weights medium and large-scale models. Our experiments reveal substantial performance gaps, underscoring the limitations of current VLMs in culturally diverse contexts. By centering cultural diversity and multilinguality, IndicVisionBench establishes a reproducible evaluation framework that paves the way for more inclusive multimodal research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。