构建首个吴语机器翻译评估基准,助力濒危语言技术保护
Machine Translation Evaluation Benchmark for Wu Chinese: Workflow and Analysis
- 人工翻译构建开源吴语数据集,支持中英互译评测
- 覆盖吴语特有表达,验证现有模型在低资源语言表现
- 适合研究低资源语言、方言保护与自然语言处理的学者
我们引入FLORES+数据集作为现代吴语机器翻译的评估基准,并展示其与现有吴语数据的兼容性。吴语与普通话、粤语等其他汉语方言互不互通,但使用高度重叠的汉字系统。吴语使用者人口数量为中国第二大语言群体,但年轻一代使用率显著下降。本文将吴语定位为文本层面的低资源语言,针对其机器翻译模型面临挑战提出解决方案。贡献包括:(1)一个开源的人工翻译数据集;(2)完整的数据集创建与验证实验文档;(3)初步的吴语标准化与分词工具;(4)本数据集的优势与局限性分析,以及对其他低资源语言的启示。
原文摘要 · Abstract (English)
We introduce a FLORES+ dataset as an evaluation benchmark for modern Wu Chinese machine translation models and showcase its compatibility with existing Wu data. Wu Chinese is mutually unintelligible with other Sinitic languages such as Mandarin and Yue (Cantonese), but uses a set of Hanzi (Chinese characters) that profoundly overlaps with others. The population of Wu speakers is the second largest among languages in China, but the language has been suffering from significant drop in usage especially among the younger generations. We identify Wu Chinese as a textually low-resource language and address challenges for its machine translation models. Our contributions include: (1) an open-source, manually translated dataset, (2) full documentations on the process of dataset creation and validation experiments, (3) preliminary tools for Wu Chinese normalization and segmentation, and (4) benefits and limitations of our dataset, as well as implications to other low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。