让视觉语言模型不再只说英语,支持多语言输出且不丢视觉理解能力。
Breaking Language Barriers in Visual Language Models via Multilingual Textual Regularization
- 用纯文本多语言数据在视觉指令微调中持续注入,保留语言模型原有多语能力。
- 多语言下语言保真度显著提升,视觉任务表现无下降,实现双优。
- 适合希望模型支持全球语言的开发者,尤其对非英语用户重要。
视觉语言模型(VLMs)虽在多模态理解上进展迅速,但常仅生成英语回答,无论输入为何种语言,这种现象称为图像诱导保真度损失(IFL),源于多模态多语言训练数据不足。为此,我们提出一种连续多语言融合策略,在视觉指令微调阶段注入纯文本多语言数据,以保留语言模型原有的多语能力。大量实验表明,该方法显著提升跨语言的语言保真度,且不损害视觉性能。我们还探索了模型合并,虽能提升语言保真度,但会牺牲视觉表现。相比之下,我们的核心方法实现了稳健的多语言对齐,无性能折损,为全球VLM应用提供了可扩展、高效缓解IFL的路径。
原文摘要 · Abstract (English)
Rapid advancements in Visual Language Models (VLMs) have transformed multimodal understanding but are often constrained by generating English responses regardless of the input language. This phenomenon has been termed as Image-induced Fidelity Loss (IFL) and stems from limited multimodal multilingual training data. To address this, we propose a continuous multilingual integration strategy that injects text-only multilingual data during visual instruction tuning, preserving the language model's original multilingual capabilities. Extensive evaluations demonstrate that our approach significantly improves linguistic fidelity across languages without degradation in visual performance. We also explore model merging, which improves language fidelity but comes at the cost of visual performance. In contrast, our core method achieves robust multilingual alignment without trade-offs, offering a scalable and effective path to mitigating IFL for global VLM adoption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。