构建首个英译瑞语翻译腔对比数据集,揭示模型偏好机械翻译
A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
- 设计英译瑞语对比数据集,包含翻译腔与自然表达对
- 小规模模型多选翻译腔,即使无原文也倾向机械表达
- 适合研究语言模型生成自然度、非英语翻译优化
翻译常带有源语言痕迹,称为翻译腔。本文提出首个免费公开的英-瑞典语数据集,对比翻译腔句式与自然表达,用于探测语言模型内在偏好。数据集包含错误标签及问题描述。实验表明,小型瑞典语和多语言大模型普遍偏好翻译腔表达。当移除英文原文时,人类自然表达更受青睐,说明源文暴露会诱导模型选择字面翻译;但即便无上下文,模型仍常倾向翻译腔。该数据集与发现为提升非英语语言自然输出提供了资源与基准。
原文摘要 · Abstract (English)
Translations often carry traces of the source language, a phenomenon known as translationese. We introduce the first freely available English-to-Swedish dataset contrasting translationese sentences with idiomatic alternatives, designed to probe intrinsic preferences of language models. It includes error tags and descriptions of the problems in the original translations. In experiments evaluating smaller Swedish and multilingual LLMs with our dataset, we find that they often favor the translationese phrasing. Human alternatives are chosen more often when the English source sentence is omitted, indicating that exposure to the source biases models toward literal translations, although even without context models often prefer the translationese variant. Our dataset and findings provide a resource and benchmark for developing models that produce more natural, idiomatic output in non-English languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。