探究多语言大模型生成文本是否带有翻译腔,揭示其与人工翻译的异同。
An Investigation of Translationese in the Generations of Multilingual Large Language Models
- 用翻译特征指标评估五种语言的生成文本。
- 发现多语言大模型生成文本具明显翻译腔特征。
- 适合关注生成文本自然度与语言风格的研究者参考。
跨语言翻译文本常带有特定痕迹,称为‘翻译腔’(translationese)。多语言大语言模型(MLLMs)可生成多种语言文本,但尚不清楚其生成内容是否类似从英语或其他语言内部翻译而来,从而产生翻译腔。本文提出两个问题:(1) MLLMs生成的文本是否具有翻译腔?(2) MLLMs产生的翻译腔与人工直接翻译有何不同?我们采用已确立的翻译文本识别指标,评估最先进的多语言大模型在五种语言下的生成文本,并与非翻译文本及人工写作基准进行对比,以分离翻译腔与其他干扰因素。通过高精度分类模型、逐项语言特征方差分析,以及德语和西班牙语子集的人工标注,我们评估了MLLM生成文本中的翻译腔含量,并识别出区别于典型翻译干扰的关键特征。
原文摘要 · Abstract (English)
Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs' generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。