arXiv:2609.06970cs.CL2026-09

打造首个粤语推理大模型,解决低资源语言训练难题。

CantoneseLLM v2: Reasoning in a Low-Resource Language

论文配图:CantoneseLLM v2: Reasoning in a Low-Resource Language
图 1 · 摘自论文原文
  • 基于Qwen3构建,融合7.84亿粤语及香港相关语料进行多阶段训练
  • 30B-A3B模型在HKCanto-Eval达73.16分,接近原基线性能
  • 首次实现粤语推理行为保留,适合粤语研究与本地化应用

粤语虽广泛使用,但书面数据匮乏,缺乏可用于模型训练的大型粤语推理语料。我们开发并发布CantoneseLLM v2,基于Qwen3 8B和30B-A3B模型,通过上下文预训练(CPT)在7.84亿粤语及香港相关语料上训练,结合聊天向量融合、监督微调(SFT)、DPO与强化学习视觉回复(RLVR)。评估显示,聊天向量融合保留了原始模型的推理语言特性,而有限粤语推理数据下的SFT会显著缩短或删除推理链,降低基准表现。DPO恢复部分推理结构,尤其对8B模型有效;而引入粤语与繁体中文双重约束的RLVR训练,成功实现语言对齐并完全恢复性能。30B-A3B模型在HKCanto-Eval上达到73.16分,仅比其融合检查点低1.20分,且保持了缺失的粤语推理行为。我们公开模型检查点、训练环境及十三年繁体中文Common Crawl数据集,模型可访问于https://huggingface.co/collections/hon9kon9ize/cantonesellm-v20。

原文摘要 · Abstract (English)

Cantonese is widely spoken but remains low-resource in written data, with no large corpus of native Cantonese reasoning traces available for model training. We develop and release CantoneseLLM v2, comprising models based on Qwen3 8B and 30B-A3B. The models are trained through CPT on 784 million Cantonese and Hong Kong-related tokens, chat-vector merging, SFT, DPO, and RLVR. Evaluation across the training stages shows that chat-vector merging transfers instruction following but preserves the donor model's reasoning language, while SFT with limited Cantonese reasoning data substantially shortens or removes reasoning traces and reduces benchmark performance. DPO restores the reasoning-block format, particularly for the 8B model, but recovers only part of the lost performance. The RLVR training with Cantonese language and Traditional Chinese scripts as multiplicative constraints introduced Cantonese language alignment and restored the lost performance. The 30B-A3B model reaches 73.16 on HKCanto-Eval, within 1.20 points of its merged checkpoint, while retaining the Cantonese reasoning behaviour absent from that checkpoint. We release the model checkpoints, the training environments, and a thirteen-year Traditional Chinese Common Crawl dataset. The models can be accessed at https://huggingface.co/collections/hon9kon9ize/cantonesellm-v20

粤语模型低资源语言推理生成大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。