用少量巴斯克语多模态数据,就能训练出高性能模型。
Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
- 用巴斯克语图文数据混合训练,仅20%占比就有效果。
- 无需专用巴斯克语大模型,通用模型也能表现良好。
- 开源资源助力其他低资源语言的多模态模型开发。
当前多模态大语言模型在多个高难度任务上表现优异。尽管商业 MLLM 在低资源语言中表现尚可,但开放科学界尚未达到类似效果。本文以巴斯克语为例,致力于构建一个强大的多模态大语言模型。为此,我们自主构建了训练与评估用的图像-文本数据集。采用 Llama-3.1-Instruct 和专为巴斯克语优化的 Latxa 模型作为骨干网络,探索多种数据混合策略。结果表明:一、仅需约 20% 的巴斯克语多模态数据即可在巴斯克语基准测试中获得稳健性能;二、出乎意料的是,无需使用巴斯克语指令微调的骨干模型,也能实现强性能。研究结果为其他低资源语言的 MLLM 开发提供了路径,并公开发布了全部资源。
原文摘要 · Abstract (English)
Current Multimodal Large Language Models exhibit very strong performance for several demanding tasks. While commercial MLLMs deliver acceptable performance in low-resource languages, comparable results remain unattained within the open science community. In this paper, we aim to develop a strong MLLM for a low-resource language, namely Basque. For that purpose, we develop our own training and evaluation image-text datasets. Using two different Large Language Models as backbones, the Llama-3.1-Instruct model and a Basque-adapted variant called Latxa, we explore several data mixtures for training. We show that: i) low ratios of Basque multimodal data (around 20%) are already enough to obtain solid results on Basque benchmarks, and ii) contrary to expected, a Basque instructed backbone LLM is not required to obtain a strong MLLM in Basque. Our results pave the way to develop MLLMs for other low-resource languages by openly releasing our resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。