构建多语言视觉语言模型训练与评估资源,支持五种欧洲语言。
Multilingual Training and Evaluation Resources for Vision-Language Models

- 通过合成生成+人工标注的混合方法构建跨语言数据集。
- 多语言训练使非英语评测表现显著提升,英语任务也受益。
- 适合关注多语言AI、跨语言视觉理解的研究者使用。
视觉语言模型(VLMs)近年取得快速发展,但其开发仍严重依赖英语,存在两大局限:(i) 缺乏多语言多模态训练数据集;(ii) 跨语言全面评估基准稀缺。本文提出一套涵盖英、法、德、意、西五种欧洲语言的综合性资源,采用再生-翻译范式,结合受许可限制的模型生成与人工标注,构建了训练语料Multi-PixMo,源自PixMo-Cap、PixMo-AskModelAnything和CoSyn-400k等现有数据集。在评估方面,将广泛使用的英文数据集(MMbench、ScienceQA、MME、POPE、AI2D)翻译为多语言版本。通过定性与定量的人类评估及标注者一致性分析验证资源质量。消融实验表明,相比仅用英语训练,加入多语言多模态数据能持续提升非英语任务表现,并带来对英语任务的正向迁移。三类模型实验均验证该效果。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) achieved rapid progress in the recent years. However, despite their growth, VLMs development is heavily grounded on English, leading to two main limitations: (i) the lack of multilingual and multimodal datasets for training, and (ii) the scarcity of comprehensive evaluation benchmarks across languages. In this work, we address these gaps by introducing a new comprehensive suite of resources for VLMs training and evaluation spanning five European languages (English, French, German, Italian, and Spanish). We adopt a regeneration-translation paradigm that produces high-quality cross-lingual resources by combining curated synthetic generation and manual annotation. Specifically, we build Multi-PixMo, a training corpus obtained regenerating examples from Pixmo pre-existing datasets with permissively licensed models: PixMo-Cap, PixMo-AskModelAnything, and CoSyn-400k. On the evaluation side, we construct a set of multilingual benchmarks derived translating widely used English datasets (MMbench, ScienceQA, MME, POPE, AI2D). We assess the quality of these resources through qualitative and quantitative human analyses, measuring inter-annotator agreement. Additionally, we perform ablation studies to demonstrate the impact of multilingual data, with respect to English only, in VLMs training. Experiments, comprising 3 different models show that using multilingual, multimodal examples for training VLMs aids is consistently beneficial on non-English benchmarks, with positive transfer to English as well.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。