将英语视觉语言模型适配法语,提升多语言AI可及性
Multilingual VLM Training: Adapting an English-Trained VLM to French
- 用翻译管道、LoRA微调和分阶段微调对比法
- 翻译数据质量是多语言性能的主要瓶颈
- 适合关注多语言AI落地的研究者与开发者
近年来人工智能在视觉-语言模型(VLMs)方面取得显著进展,能理解图像与文本数据。然而这些成果主要局限于英语,限制了非英语使用者的访问。本文探讨将英语训练的VLM适配至其他语言的挑战。比较了基于翻译的流程、LoRA微调以及将视觉与语言适应分离的两阶段微调策略。通过将标准多模态基准数据集翻译至目标语言,并结合母语专家的人工评估来验证效果。结果表明,数据集翻译仍是多语言VLM性能的关键瓶颈,数据质量制约了训练与评估的有效性。研究建议未来应聚焦原生语言数据集构建与翻译策略优化。
原文摘要 · Abstract (English)
Artificial intelligence has made great progress in recent years, particularly in the development of Vision--Language Models (VLMs) that understand both visual and textual data. However, these advancements remain largely limited to English, reducing their accessibility for non--English speakers. It is essential to extend these capabilities to a broader range of languages. This paper explores the challenges of adapting an English-trained VLM to different languages. To this end, we will explore and compare different methods for their performance and computational cost. We consider a translation-based pipeline, LoRA finetuning, and a two-stage finetuning strategy that separates vision adaptation from language adaptation. To evaluate these methods, we use a combination of standard multimodal benchmarks translated into the target language and manual assessments by native experts. The results reveal that dataset translation remains a major bottleneck in multilingual VLM performance, with data quality limiting the effectiveness of training and evaluation. These findings suggest that future efforts should focus on native-language dataset collection and improved translation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。