arXiv:2512.14926cs.CLcs.AI2025-12

为罗马尼亚语构建多模态数据集,用高效微调提升视觉问答与图文生成能力

Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models

  • 用开源大模型翻译并扩展Flickr30K,构建罗马尼亚语多模态数据集
  • 采用LoRA微调,70亿参数Qwen2-VL-RoVQA在视觉问答和图文生成上分别提升2.29%和4.45%
  • 显著减少语法错误,适合低资源语言AI研究者使用

关注低资源语言是实现生成式AI普惠化的重要一步。本文致力于缩小罗马尼亚语的多模态自然语言处理资源差距。我们将广泛使用的Flickr30K数据集翻译为罗马尼亚语,并利用开源大语言模型进一步扩展用于视觉问答任务。通过在罗马尼亚语视觉问答上微调开源多模态模型(VLMs),验证了新数据集的有效性。选用三种主流模型家族:LLaMA 3.2、LLaVA 1.6 和 Qwen2。微调采用参数高效的LoRA方法。实验表明,所训练模型在罗马尼亚语视觉问答任务及未训练过的图像描述生成任务中均表现更优。其中,70亿参数的Qwen2-VL-RoVQA在两项任务上分别较原版提升2.29%和4.45%的BERTScore F1值。此外,模型在语法错误方面显著减少,表明其不仅提升了语言理解,也增强了罗马尼亚语表达的流畅性。

原文摘要 · Abstract (English)

Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We translate the widely known Flickr30K dataset into Romanian and further extend it for visual question answering by leveraging open-source LLMs. We demonstrate the usefulness of our datasets by fine-tuning open-source VLMs on Romanian visual question answering. We select VLMs from three widely used model families: LLaMA 3.2, LLaVA 1.6, and Qwen2. For fine-tuning, we employ the parameter-efficient LoRA method. Our models show improved Romanian capabilities in visual QA, as well as on tasks they were not trained on, such as Romanian image description generation. The seven-billion-parameter Qwen2-VL-RoVQA obtains top scores on both tasks, with improvements of +2.29% and +4.45% in BERTScore F1 on VQA and captioning, respectively, over its original version. Finally, the models show substantial reductions in grammatical errors compared to their original forms, indicating improvements not only in language understanding but also in Romanian fluency.

多模态低资源语言指令微调LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。