arXiv:2605.31401cs.CL2026-05

为罗马尼亚语构建视觉语言模型,提升低资源语言的多模态理解能力

"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models

论文配图:"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
图 1 · 摘自论文原文
  • 翻译英文多模态数据集至罗马尼亚语,保持视觉对齐
  • 适配罗马尼亚语的模型在多个任务上超越更大规模模型
  • 构建本土化评测集HoraVQA,覆盖日常文化场景

视觉语言模型(VLMs)主要沿袭纯文本大模型路径,在英语基准上表现优异,但在低资源语言上性能急剧下降,因缺乏大规模图文语料库和文化相关的评估体系。本文系统研究了为罗马尼亚语构建专用VLM的全流程,涵盖从数据构建到架构选择。将现有的英文VLM训练与评估语料库翻译为罗马尼亚语,通过机器翻译处理文本注释及图像内文字,保留视觉对齐的同时适配语言内容。基于此数据,训练并消融一系列VLM,分析(i)不同规模与预训练方式的视觉主干网络、(ii)从多语言到罗马尼亚语适配的文本主干模型、(iii)OCR风格图文数据的贡献。进一步构建了基于罗马尼亚日常场景的文化原生评测集HoraVQA。结果表明,罗马尼亚语适配的VLM持续优于同尺寸基线模型,并在所有评估基准上超越下一量级更大的模型。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languages, where neither large-scale image-text corpora nor culturally grounded evaluations exist. We present a systematic study of building a language-specific VLM for Romanian, covering the full pipeline from data construction to architectural choices. We translate established English VLM training and evaluation corpora into Romanian, applying machine translation to textual annotations and to in-image text, preserving visual grounding while adapting the textual content. Using this data, we train and ablate a series of VLMs to isolate the contribution of (i) vision backbones of varying scale and pretraining, (ii) language backbones from multilingual to Romanian-adapted LLMs, and (iii) OCR-style image-text data. We further curate HoraVQA, a culturally native evaluation set grounded in Romanian everyday scenes. Romanian-adapted VLMs consistently outperform their same-sized counterparts and, across all evaluated benchmarks, even surpass models from the next larger size category.

视觉语言模型低资源语言罗马尼亚语多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。