arXiv:2502.12852cs.CL2025-02ACL被引 6

构建205种语言的跨模态主题匹配基准,揭示大模型在低资源语言上表现严重不足。

MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

  • 构建覆盖205种语言的多模态评测集,涵盖跨模态与纯文本主题匹配。
  • 低资源语言如N'Koo上模型表现仅达随机水平,跨模态能力远弱于文本理解。
  • 开源大模型无法有效利用多图像描述,提示其多图任务处理仍不成熟。

现有多语言视觉-语言(VL)评测集通常仅覆盖少数语言,导致大型视觉-语言模型(LVLMs)评估主要集中在高资源语言。为解决此问题,我们提出MVL-SIB,一个大规模多语言视觉-语言基准,涵盖205种语言,用于评估跨模态与纯文本主题匹配,远超现有最多样化基准的100余种语言。我们在MVL-SIB上评估了一系列开源LVLMs及GPT-4o(-mini)。结果表明,LVLMs在低资源语言的跨模态主题匹配中表现不佳,在如N'Koo等语言上性能接近随机水平。分析显示,随着语言资源减少,模型的跨模态支持能力下降幅度显著超过文本支持能力。此外,开源模型未因使用多张图像表示同一主题而提升性能,表明其尚未充分掌握多图像任务。通过与其它多语言VL基准的相关性分析,我们证明MVL-SIB是全面探测LVLM多语言理解能力的有效工具。

原文摘要 · Abstract (English)

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Consequently, evaluations of large vision-language models (LVLMs) predominantly target high-resource languages, underscoring the need for evaluation data for low-resource languages. To address this limitation, we introduce MVL-SIB, a massively multilingual vision-language benchmark that evaluates both cross-modal and text-only topical matching across 205 languages -- over 100 more than the most multilingual existing VL benchmarks encompass. We then benchmark a range of of open-weight LVLMs together with GPT-4o(-mini) on MVL-SIB. Our results reveal that LVLMs struggle in cross-modal topic matching in lower-resource languages, performing no better than chance on languages like N'Koo. Our analysis further reveals that VL support in LVLMs declines disproportionately relative to textual support for lower-resource languages, as evidenced by comparison of cross-modal and text-only topical matching performance. We further observe that open-weight LVLMs do not benefit from representing a topic with more than one image, suggesting that these models are not yet fully effective at handling multi-image tasks. By correlating performance on MVL-SIB with other multilingual VL benchmarks, we highlight that MVL-SIB serves as a comprehensive probe of multilingual VL understanding in LVLMs.

多语言跨模态评测基准低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。