对比大模型与神经网络翻译的机器翻译语体特征
Decoding Machine Translationese in English-Chinese News: LLMs vs. NMTs
- 构建四类新闻语料库,用五层特征识别机器翻译特有模式
- 两类模型均存在翻译腔,大模型词汇更丰富但中文句短、连词多
- 国产与国外大模型翻译风格无显著差异,适合内容检测与评估
本研究探讨机器翻译语体(MTese)在英文到中文新闻文本中的表现,聚焦此前研究较少的语对。我们构建了一个包含4个子语料库的大规模数据集,并采用全面的五层特征集。通过卡方排序算法进行特征选择,应用于分类与聚类任务。研究结果证实,神经机器翻译系统(NMTs)和大型语言模型(LLMs)均存在机器翻译语体。原始中文文本与两类模型输出几乎可完美区分。机器翻译输出的典型特征包括更短的句子长度和更高频的转折性连词使用。在对比中,模型分类准确率约70%,其中大模型表现出更高的词汇多样性,而神经网络模型更多使用括号。此外,专用翻译大模型相比通用模型词汇多样性更低但因果连词使用更高。最后,中国公司开发的大模型与外国同类模型之间无显著差异。
原文摘要 · Abstract (English)
This study explores Machine Translationese (MTese) -- the linguistic peculiarities of machine translation outputs -- focusing on the under-researched English-to-Chinese language pair in news texts. We construct a large dataset consisting of 4 sub-corpora and employ a comprehensive five-layer feature set. Then, a chi-square ranking algorithm is applied for feature selection in both classification and clustering tasks. Our findings confirm the presence of MTese in both Neural Machine Translation systems (NMTs) and Large Language Models (LLMs). Original Chinese texts are nearly perfectly distinguishable from both LLM and NMT outputs. Notable linguistic patterns in MT outputs are shorter sentence lengths and increased use of adversative conjunctions. Comparing LLMs and NMTs, we achieve approximately 70% classification accuracy, with LLMs exhibiting greater lexical diversity and NMTs using more brackets. Additionally, translation-specific LLMs show lower lexical diversity but higher usage of causal conjunctions compared to generic LLMs. Lastly, we find no significant differences between LLMs developed by Chinese firms and their foreign counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。