arXiv:2602.14073cs.CLcs.AI2026-02Conference of the …被引 3

用自动化翻译+少量人工,让视觉语言模型学会波兰语。

Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Framework

  • 用自动翻译和过滤现有数据,辅以合成波兰语数据训练模型。
  • 在波兰语评测中比原模型高9.5分,生成句子更符合语法。
  • 适合想低成本适配低资源语言的多模态研究者。

大多数视觉语言模型基于英语数据训练,限制了其在其他语言和文化背景下的表现,影响非英语用户使用并阻碍多元语言与文化系统的开发。本文复现并改进LLaVA-Next方法,构建一系列波兰语视觉语言模型。通过全自动流程翻译与筛选多模态数据,并补充用于OCR和文化特定任务的合成波兰语数据。尽管训练数据几乎全靠自动翻译且人工干预极少,模型仍表现优异:在波兰语适配的MMBench上相较LLaVA-1.6-Vicuna-13B提升9.5%,生成内容经人工评估在语言正确性上更优。结果表明,大规模自动翻译结合轻量级过滤可有效构建低资源语言的高质量多模态模型。部分挑战仍存,如文化覆盖不足与评估局限。为促进后续研究,我们公开发布模型及评测数据集。

原文摘要 · Abstract (English)

Most vision-language models (VLMs) are trained on English-centric data, limiting their performance in other languages and cultural contexts. This restricts their usability for non-English-speaking users and hinders the development of multimodal systems that reflect diverse linguistic and cultural realities. In this work, we reproduce and adapt the LLaVA-Next methodology to create a set of Polish VLMs. We rely on a fully automated pipeline for translating and filtering existing multimodal datasets, and complement this with synthetic Polish data for OCR and culturally specific tasks. Despite relying almost entirely on automatic translation and minimal manual intervention to the training data, our approach yields strong results: we observe a +9.5% improvement over LLaVA-1.6-Vicuna-13B on a Polish-adapted MMBench, along with higher-quality captions in generative evaluations, as measured by human annotators in terms of linguistic correctness. These findings highlight that large-scale automated translation, combined with lightweight filtering, can effectively bootstrap high-quality multimodal models for low-resource languages. Some challenges remain, particularly in cultural coverage and evaluation. To facilitate further research, we make our models and evaluation dataset publicly available.

视觉语言模型多语言波兰语低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。