打造多语言视觉语言模型,提升跨语言理解与文化情境识别能力。
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
- 基于Tower+构建多语言文本编码器,融合视觉与文化上下文微调
- 在ALM-Bench、Multi30K等任务上超越更大规模模型,尤其擅长跨文化任务
- 开源模型、数据集和训练方法,推动多语言多模态研究
尽管视觉语言模型(VLMs)取得显著进展,但多数工作仍采用以英语为中心的设计,限制了其在多语言场景下的表现。本文通过系统实证研究,分析了训练数据构成、编码器选择和文本骨干网络等多语言设计因素的影响。基于此,我们提出TowerVision,一个面向图像-文本与视频-文本任务的开源多语言VLM家族,建立在多语言纯文本模型Tower+之上。TowerVision在多个多模态多语言基准上表现优异,尤其在文化语境相关任务和多模态翻译中展现出突出能力。通过在微调阶段引入视觉与文化上下文,我们的模型在ALM-Bench和Multi30K(图像任务)以及ViMUL-Bench(视频任务)上超越了使用更大规模数据训练的现有方法。同时,我们发布了高质量、精心筛选的视觉语言数据集VisionBlocks。研究发现,多语言多模态训练数据能显著提升跨语言泛化能力,既支持高资源到低资源语言的迁移,也支持反向迁移;且指令微调的大语言模型并非总是最优初始化点。为促进后续研究,所有模型、数据及训练方案均已公开发布。
原文摘要 · Abstract (English)
Despite significant advances in vision-language models (VLMs), most existing work follows an English-centric design process, limiting their effectiveness in multilingual settings. In this work, we provide a comprehensive empirical study analyzing the impact of several multilingual design choices, such as training data composition, encoder selection, and text backbones. The result is TowerVision, a family of open multilingual VLMs for both image-text and video-text tasks, built upon the multilingual text-only model Tower+. TowerVision achieves competitive performance on multiple multimodal multilingual benchmarks and shows particular strength in culturally grounded tasks and multimodal translation. By incorporating visual and cultural context during fine-tuning, our models surpass existing approaches trained on substantially larger datasets, as demonstrated on ALM-Bench and Multi30K (image tasks) and ViMUL-Bench (video tasks). Alongside the models, we release VisionBlocks, a high-quality, curated vision-language dataset. Our findings highlight that multilingual vision-language training data substantially improves cross-lingual generalization -- both from high-resource to underrepresented languages and vice versa -- and that instruction-tuned LLMs are not always the optimal initialization point. To support further research, we publicly release all models, data, and training recipes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。