为爱尔兰语打造双语大模型,解决低资源语言发展难题
Qomhra: A Bilingual Irish and English Large Language Model
- 用双语持续预训练+指令微调构建爱尔兰语大模型
- 翻译海量英文数据并合成首个爱尔兰语偏好数据集
- 性能比现有开源模型提升最高44%,适合濒危语言研究者
大型语言模型研究长期集中于主流语言,导致爱尔兰语等低资源语言被忽视。本文提出Qomhrá,一个在极低资源条件下构建的双语爱尔兰语与英语大模型。完整流程包括双语持续预训练、指令微调,以及通过提示大模型生成“接受”与“拒绝”响应来合成人类偏好数据,验证结果显示其符合母语爱尔兰语使用者的判断。我们评估了多个闭源大模型在爱尔兰语生成中的表现,发现Gemini-2.5-Pro在母语和非母语爱尔兰语使用者中评分最高,与大模型作为裁判的评价存在差异,表明当前模型与爱尔兰语社区存在不匹配。随后,我们利用Gemini-2.5-Pro将大规模英文指令微调数据集翻译为爱尔兰语,并合成首个爱尔兰语人类偏好数据集。我们在多项基准上全面评估Qomhrá,涵盖翻译、性别理解、主题识别和世界知识;结果表明,在爱尔兰语上性能提升达29%,英语上提升达44%,显著优于现有开源爱尔兰语基线模型UCCIX。该框架为爱尔兰语及其他低资源语言的大模型开发提供了重要参考。
原文摘要 · Abstract (English)
Large language model (LLM) research and development has overwhelmingly focused on the world's major languages, leading to under-representation of low-resource languages such as Irish. This paper introduces \textbf{Qomhrá}, a bilingual Irish and English LLM, developed under extremely low-resource constraints. A complete pipeline is outlined spanning bilingual continued pre-training, instruction tuning, and the synthesis of human preference data for future alignment training. We focus on the lack of scalable methods to create human preference data by proposing a novel method to synthesise such data by prompting an LLM to generate ``accepted'' and ``rejected'' responses, which we validate as aligning with L1 Irish speakers. To select an LLM for synthesis, we evaluate the top closed-weight LLMs for Irish language generation performance. Gemini-2.5-Pro is ranked highest by L1 and L2 Irish-speakers, diverging from LLM-as-a-judge ratings, indicating a misalignment between current LLMs and the Irish-language community. Subsequently, we leverage Gemini-2.5-Pro to translate a large scale English-language instruction tuning dataset to Irish and to synthesise a first-of-its-kind Irish-language human preference dataset. We comprehensively evaluate Qomhrá across several benchmarks, testing translation, gender understanding, topic identification, and world knowledge; these evaluations show gains of up to 29\% in Irish and 44\% in English compared to the existing open-source Irish LLM baseline, UCCIX. The results of our framework provide insight and guidance to developing LLMs for both Irish and other low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。