探究汉字拆分如何提升中文多词表达翻译效果
An Empirical Study on Chinese Character Decomposition in Multiword Expression-Aware Neural Machine Translation
- 将汉字拆分为部件,增强模型对中文词汇语义的理解
- 实验证明拆分能有效缓解多词表达的翻译歧义问题
- 适合关注中文机器翻译与语义表示的研究者
词义、表征与理解在自然语言理解、处理与生成任务中起核心作用。许多任务的困难源于多词表达(MWEs),其带来的歧义、习语性、罕见使用及丰富变体增加了处理难度。西方语言如英语在该领域已有显著进展,得益于成熟研究生态与充足计算资源。但中文等东亚语言在此方面仍滞后。虽子词建模(如BPE)在西方语言中成功缓解生僻词问题,但难以直接应用于汉字这类表意文字。本文系统研究了在面向多词表达的神经机器翻译中,汉字拆分技术的作用。通过实验分析汉字拆分如何更好地保留原词原义,并有效应对多词表达的翻译挑战。
原文摘要 · Abstract (English)
Word meaning, representation, and interpretation play fundamental roles in natural language understanding (NLU), natural language processing (NLP), and natural language generation (NLG) tasks. Many of the inherent difficulties in these tasks stem from Multi-word Expressions (MWEs), which complicate the tasks by introducing ambiguity, idiomatic expressions, infrequent usage, and a wide range of variations. Significant effort and substantial progress have been made in addressing the challenging nature of MWEs in Western languages, particularly English. This progress is attributed in part to the well-established research communities and the abundant availability of computational resources. However, the same level of progress is not true for language families such as Chinese and closely related Asian languages, which continue to lag behind in this regard. While sub-word modelling has been successfully applied to many Western languages to address rare words improving phrase comprehension, and enhancing machine translation (MT) through techniques like byte-pair encoding (BPE), it cannot be applied directly to ideograph language scripts like Chinese. In this work, we conduct a systematic study of the Chinese character decomposition technology in the context of MWE-aware neural machine translation (NMT). Furthermore, we report experiments to examine how Chinese character decomposition technology contributes to the representation of the original meanings of Chinese words and characters, and how it can effectively address the challenges of translating MWEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。