arXiv:2503.03702cs.CL2025-03EMNLP被引 1

构建超20亿词元粤语数据集,提升大模型粤语能力

Developing and Utilizing a Large-Scale Cantonese Dataset for Multi-Tasking in Large Language Models

  • 从多源数据收集并清洗,构建超20亿词元高质量粤语语料
  • 在4个粤语基准上达到当前最优(SOTA)性能
  • 不仅提升粤语任务表现,还增强主流语言任务能力

高质量数据资源对大语言模型训练至关重要,尤其对于粤语等低资源语言。尽管粤语母语者超过8500万,但因普通话主导、社群分散、编码与输入方式多样,以及海外使用者偏好英语,导致其在自然语言处理领域仍属低资源。此外,粤语丰富的口语词汇、英文借词及语码转换特征增加了语料采集与处理难度。为此,我们从开源语料、香港特定论坛、维基百科和Common Crawl等多源获取粤语文本,通过语言过滤、质量筛选、内容清理和去重等严格处理步骤,成功构建了超过20亿词元的高质量粤语语料库,用于大模型训练。进一步在精选粤语任务上进行监督微调(SFT),显著提升模型在特定应用中的表现。模型在完成训练后,在4个粤语基准测试中均达到当前最优(SOTA)水平。此外,该模型在其他主流语言任务上也展现出性能提升。

原文摘要 · Abstract (English)

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a low-resource language in the field of natural language processing (NLP) due to factors such as the dominance of Mandarin, lack of cohesion within the Cantonese-speaking community, diversity in character encoding and input methods, and the tendency of overseas Cantonese speakers to prefer using English. In addition, rich colloquial vocabulary of Cantonese, English loanwords, and code-switching characteristics add to the complexity of corpus collection and processing. To address these challenges, we collect Cantonese texts from a variety of sources, including open source corpora, Hong Kong-specific forums, Wikipedia, and Common Crawl data. We conduct rigorous data processing through language filtering, quality filtering, content filtering, and de-duplication steps, successfully constructing a high-quality Cantonese corpus of over 2 billion tokens for training large language models. We further refined the model through supervised fine-tuning (SFT) on curated Cantonese tasks, enhancing its ability to handle specific applications. Upon completion of the training, the model achieves state-of-the-art (SOTA) performance on four Cantonese benchmarks. After training on our dataset, the model also exhibits improved performance on other mainstream language tasks.

粤语大模型数据集多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。