构建中英被动句双向数据集,揭示机器翻译对被动语态的处理差异。
Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation
- 从五个双语语料库提取并自动标注7万多组被动句对。
- 模型更依赖源文语态,中译英时被动语态保留率更高。
- 商业模型得分高但创意不足,大模型更擅用多样表达。
机器翻译评估已超越传统指标,转向关注具体语言现象。针对中英语言对,被动句的构建与分布因语言差异而不同,需在翻译中特别关注。本文提出一个双向多领域被动句数据集,源自五个中英平行语料库,通过人工翻译标准进行结构标签的自动化标注,并建立经人工验证的测试集。数据集包含73,965组平行句子对(2,358,731个英文单词,3,498,229个中文字符)。我们使用该数据集评估了两种先进的开源MT系统,以及四种商用模型。结果表明,与人类不同,模型更受源文本语态影响,倾向于在双向翻译中保持被动语态。然而,模型对中文被动句低频且多负面语境的认知,使其在英译中时语态一致性优于中译英。商用NMT模型在指标上得分更高,但大语言模型展现出更强的替代翻译多样性。数据集与标注脚本可应要求提供。
原文摘要 · Abstract (English)
Machine Translation (MT) evaluation has gone beyond metrics, towards more specific linguistic phenomena. Regarding English-Chinese language pairs, passive sentences are constructed and distributed differently due to language variation, thus need special attention in MT. This paper proposes a bidirectional multi-domain dataset of passive sentences, extracted from five Chinese-English parallel corpora and annotated automatically with structure labels according to human translation, and a test set with manually verified annotation. The dataset consists of 73,965 parallel sentence pairs (2,358,731 English words, 3,498,229 Chinese characters). We evaluate two state-of-the-art open-source MT systems with our dataset, and four commercial models with the test set. The results show that, unlike humans, models are more influenced by the voice of the source text rather than the general voice usage of the source language, and therefore tend to maintain the passive voice when translating a passive in either direction. However, models demonstrate some knowledge of the low frequency and predominantly negative context of Chinese passives, leading to higher voice consistency with human translators in English-to-Chinese translation than in Chinese-to-English translation. Commercial NMT models scored higher in metric evaluations, but LLMs showed a better ability to use diverse alternative translations. Datasets and annotation script will be shared upon request.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。