为埃及和摩洛哥方言打造兼顾语言与文化的中文大模型
NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
- 用本地语言和文化数据合成训练语料,避免文化偏差
- 30亿参数模型在理解与价值观对齐上超越同类小模型
- 开源方法与数据,推动低资源社区语言模型发展
提升大语言模型对低资源语言的处理能力是关键研究方向。现有方法多依赖英文语料翻译生成合成数据,虽能提升语言理解与翻译能力,但常导致模型偏向源语言文化,难以体现本地社区的文化传承与价值观。本文提出一种面向特定社区的合成与检索增强预训练数据构建方法,综合考虑语言、文化传承与价值观三要素。以埃及与摩洛哥方言为测试案例,因其语言文化丰富且当前在大模型中严重缺失。作为概念验证,我们开发了NileChat——一个30亿参数的埃及与摩洛哥阿拉伯语大模型,专为当地社区定制,融合其语言、文化与价值观。在多项理解、翻译及文化价值观对齐基准测试中,NileChat表现优于同规模现有阿拉伯语模型,并达到更大模型水平。本工作通过受控合成数据生成与检索增强预训练,实现摩洛哥达里贾与埃及阿拉伯语(含阿拉伯字母变体)的高质量建模,推动低资源社区阿拉伯语自然语言处理发展。相关方法、数据与模型已开源:https://github.com/UBC-NLP/nilechat。
原文摘要 · Abstract (English)
Enhancing the linguistic capabilities of Large Language Models (LLMs) to include low-resource languages is a critical research area. Current research directions predominantly rely on synthetic data generated by translating English corpora, which, while demonstrating promising linguistic understanding and translation abilities, often results in models aligned with source language culture. These models frequently fail to represent the cultural heritage and values of local communities. This work proposes a methodology to create both synthetic and retrieval-based pre-training data tailored to a specific community, considering its (i) language, (ii) cultural heritage, and (iii) cultural values. We demonstrate our methodology using Egyptian and Moroccan dialects as testbeds, chosen for their linguistic and cultural richness and current underrepresentation in LLMs. As a proof-of-concept, we develop NileChat, a 3B parameter Egyptian and Moroccan Arabic LLM adapted for Egyptian and Moroccan communities, incorporating their language, cultural heritage, and values. Our results on various understanding, translation, and cultural and values alignment benchmarks show that NileChat outperforms existing Arabic-aware LLMs of similar size and performs on par with larger models. This work addresses Arabic dialect in LLMs with a focus on cultural and values alignment via controlled synthetic data generation and retrieval-augmented pre-training for Moroccan Darija and Egyptian Arabic, including Arabizi variants, advancing Arabic NLP for low-resource communities. We share our methods, data, and models with the community to promote the inclusion and coverage of more diverse communities in cultural LLM development: https://github.com/UBC-NLP/nilechat .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。