arXiv:2602.16469cs.CL2026-02

不同源语言的翻译风格影响小模型学习效果,揭示翻译语体的深层影响

Dialects of Translationese Shape Language Model Learning

  • 用24种语言的机器翻译数据训练英文模型,分析源语言差异的影响
  • 词汇多样性决定整体困惑度,语法表现与源语言类型相似性强相关
  • 翻译质量是模型性能的关键预测因子,适合多语言研究者参考

机器翻译数据广泛用于多语言自然语言处理,尤其在本族语数据稀缺时。然而,翻译文本与本族语文本存在系统性差异,这种现象称为翻译语体(translationese),反映了源语言痕迹及翻译本身的特点。本文研究了在机器翻译数据上训练的小型英文语言模型,重点分析不同源语言的翻译语体如何影响语言可接受性判断和跨领域语言建模。我们使用来自24种类型学和资源分布各异的源语言的英语翻译文本进行训练,实现对源语言和语料属性影响的系统分析。结果表明:源语言显著影响模型行为——整体困惑度主要受翻译语料词汇多样性驱动,而语法表现则在数据量充足时与源语言类型学相似性高度相关。即使翻译质量本身也是语言建模范畴的重要预测因子。

原文摘要 · Abstract (English)

Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it reflects both traces of the source language and characteristic properties of translation itself. In this paper, we study how training on machine-translated data affects small English language models, focusing on how translationese from different source languages shapes linguistic acceptability judgments and language modeling for different domains. We train models on English text translated from 24 typologically and resource-diverse source languages, enabling a systematic analysis of how source language and corpus properties influence what models learn. Our results show that the source language has a clear impact on model behavior: general perplexity is more driven by the lexical diversity of the translated corpus, but grammatical performance is strongly correlated to typological similarity to English if trained on enough data. Even translation quality is a strong predictor of language modeling performance.

语言模型翻译语体多语言NLP数据偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。