arXiv:2505.17816cs.CL2025-05被引 9

解决粤语书面化翻译数据稀缺问题,提升中粤互译质量

Low-Resource NMT: A Case Study on the Written and Spoken Languages in Hong Kong

  • 从双语维基百科自动挖掘7.2万组相似句对扩充训练数据
  • 在8个测试集上6项超越百度翻译,最高提升1.8点BLEU
  • 适合需要精准处理粤语书面表达的本地化应用

香港多数居民能读写标准中文,但日常口语使用粤语。粤语可转写为汉字形成书面粤语,其词汇与语法与标准中文差异显著。随着网络交流增多,中粤自动翻译需求上升。本文构建基于Transformer的中-粤书面语机器翻译系统。鉴于中粤平行语料极度稀缺,研究重点在于扩充训练数据:除收集2.8万句来自过往语言学研究和零散网络资源外,还设计有效方法,从简体与繁体维基百科的平行文章中自动提取7.2万组语义相近句子对。实验表明,利用维基百科挖掘的高相似句对可显著提升所有测试集表现。系统在8个测试集中有6项超越百度翻译的中-粤翻译,平均提升1.8点BLEU。翻译示例显示,系统能捕捉标准中文与粤语口语间的典型语言转换特征。

原文摘要 · Abstract (English)

The majority of inhabitants in Hong Kong are able to read and write in standard Chinese but use Cantonese as the primary spoken language in daily life. Spoken Cantonese can be transcribed into Chinese characters, which constitute the so-called written Cantonese. Written Cantonese exhibits significant lexical and grammatical differences from standard written Chinese. The rise of written Cantonese is increasingly evident in the cyber world. The growing interaction between Mandarin speakers and Cantonese speakers is leading to a clear demand for automatic translation between Chinese and Cantonese. This paper describes a transformer-based neural machine translation (NMT) system for written-Chinese-to-written-Cantonese translation. Given that parallel text data of Chinese and Cantonese are extremely scarce, a major focus of this study is on the effort of preparing good amount of training data for NMT. In addition to collecting 28K parallel sentences from previous linguistic studies and scattered internet resources, we devise an effective approach to obtaining 72K parallel sentences by automatically extracting pairs of semantically similar sentences from parallel articles on Chinese Wikipedia and Cantonese Wikipedia. We show that leveraging highly similar sentence pairs mined from Wikipedia improves translation performance in all test sets. Our system outperforms Baidu Fanyi's Chinese-to-Cantonese translation on 6 out of 8 test sets in BLEU scores. Translation examples reveal that our system is able to capture important linguistic transformations between standard Chinese and spoken Cantonese.

机器翻译低资源粤语数据挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。