arXiv:2501.05122cs.CLcs.CV2025-01ACL被引 3

提出多语言视觉-语言模型训练新策略,100种语言下仍保持强英文性能

Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model

  • 通过系统实验探索多语言训练配比,发现仅需25%-50%非英语数据即可提升多语言能力
  • 在13项任务、43种语言上验证,支持同时训练100种语言且不降低英文表现
  • 强调非英语OCR数据对图文理解的关键作用,构建新基准评测该能力

现有大规模视觉-语言模型(LVLM)主要基于英语数据训练,导致对非英语输入理解差且难以生成目标语言输出。现有方法虽加入多语言数据缓解问题,但缺乏对不同语言组合如何影响性能的系统认知。本文通过覆盖13个下游任务和43种语言的多阶段实验,系统研究:(1)不损害英文性能的前提下可纳入的训练语言数量;(2)预训练与指令微调阶段最优语言分布。进一步探究如何提升多语言图文理解能力,并引入新基准。结果表明,可同时包含多达100种语言,且仅需25%-50%非英语数据即可显著提升多语言表现,同时保持强大英文性能。此外,将非英语OCR数据融入预训练和指令微调至关重要。基于以上发现,我们训练出Centurio——一个支持100种语言的LVLM,在涵盖14项任务和56种语言的评估中达到当前最优表现。

原文摘要 · Abstract (English)

Most Large Vision-Language Models (LVLMs) to date are trained predominantly on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. Existing efforts mitigate these issues by adding multilingual training data, but do so in a largely ad-hoc manner, lacking insight into how different training mixes tip the scale for different groups of languages. In this work, we present a comprehensive investigation into the training strategies for massively multilingual LVLMs. First, we conduct a series of multi-stage experiments spanning 13 downstream vision-language tasks and 43 languages, systematically examining: (1) the number of training languages that can be included without degrading English performance and (2) optimal language distributions of pre-training as well as (3) instruction-tuning data. Further, we (4) investigate how to improve multilingual text-in-image understanding, and introduce a new benchmark for the task. Surprisingly, our analysis reveals that one can (i) include as many as 100 training languages simultaneously (ii) with as little as 25-50\% of non-English data, to greatly improve multilingual performance while retaining strong English performance. We further find that (iii) including non-English OCR data in pre-training and instruction-tuning is paramount for improving multilingual text-in-image understanding. Finally, we put all our findings together and train Centurio, a 100-language LVLM, offering state-of-the-art performance in an evaluation covering 14 tasks and 56 languages.

多语言视觉语言模型训练策略图文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。