首个粤语自然语言理解基准,覆盖七类任务。
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
- 构建涵盖语法语义的七项任务粤语评测集。
- 粤语微调模型整体表现最优,语法任务中单语模型更优。
- 适合粤语NLP研究者、方言自然语言处理方向人员。
粤语虽有数百万使用者,但因政策与双言现象导致资源匮乏。为填补粤语自然语言理解评估框架的空白,我们提出 extsc{CantoNLU},一个涵盖七个任务的粤语自然语言理解基准,包括词义消歧、语言可接受性判断、语言识别、自然语言推理、情感分析、词性标注和依存句法分析。同时提供四类模型基线:未在粤语上训练的普通话模型、两种通过持续预训练从普通话模型迁移而来的粤语适配模型,以及从头训练的单语粤语模型。结果表明,粤语适配模型整体表现最佳,而单语模型在语法任务上更优;普通话模型在某些场景仍具竞争力,说明当粤语领域数据稀缺时,直接迁移可能已足够。所有数据集、代码与模型权重均已公开,以促进粤语NLP研究。
原文摘要 · Abstract (English)
Cantonese, although spoken by millions, remains under-resourced due to policy and diglossia. To address this scarcity of evaluation frameworks for Cantonese, we introduce \textsc{\textbf{CantoNLU}}, a benchmark for Cantonese natural language understanding (NLU). This novel benchmark spans seven tasks covering syntax and semantics, including word sense disambiguation, linguistic acceptability judgment, language detection, natural language inference, sentiment analysis, part-of-speech tagging, and dependency parsing. In addition to the benchmark, we provide model baseline performance across a set of models: a Mandarin model without Cantonese training, two Cantonese-adapted models obtained by continual pre-training a Mandarin model on Cantonese text, and a monolingual Cantonese model trained from scratch. Results show that Cantonese-adapted models perform best overall, while monolingual models perform better on syntactic tasks. Mandarin models remain competitive in certain settings, indicating that direct transfer may be sufficient when Cantonese domain data is scarce. We release all datasets, code, and model weights to facilitate future research in Cantonese NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。