arXiv:2506.23340cs.CL2025-06

研究发现语言结构越接近英语,翻译信息损失越少。

Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family

  • 用回译法测试GPT-4和Llama 2的多语言翻译表现
  • 数据少时,离英语越近的语言翻译质量越高
  • 语言家族、语法和字形差异是关键预测因素

大语言模型在多语言翻译中虽取得显著进展,但仍面临某些语言对的挑战,尤其在训练数据有限或与英语语言差异较大时。本研究系统考察了训练数据量、语言距离和语言家族对翻译信息损失的影响。通过回译法评估GPT-4和Llama 2,使用BLEU分数和BERT相似性指标衡量翻译质量。结果表明,训练数据量与语言距离存在显著交互作用:数据充足可缓解语言差异影响,但在低资源条件下,结构上更接近英语的语言始终表现更好。多种距离度量——字形、谱系、句法和地理距离——均能有效预测翻译性能。语言家族也具有独立影响。研究揭示了大语言模型多语言翻译中的语言学约束,强调翻译质量不仅取决于数据量,还受语言间结构与类型关系的影响。

原文摘要 · Abstract (English)

Large language models have achieved impressive progress in multilingual translation, yet they continue to face challenges with certain language pairs-particularly those with limited training data or significant linguistic divergence from English. This study systematically investigates how training data, language proximity, and language family affect information loss in multilingual translation. We evaluate two large language models, GPT-4 and Llama 2, by performing round-trip translations. Translation quality was assessed using BLEU scores and BERT similarity metrics. Our results reveal a robust interaction between training data size and language distance: while abundant training data can mitigate the effects of linguistic divergence, languages structurally closer to English consistently yield higher translation quality in low-resource conditions. Among various distance metrics, orthographic, phylogenetic, syntactic, and geographical distances emerge as strong predictors of translation performance. Language family also exerts an independent influence. These findings contribute to a deeper understanding of the linguistic constraints shaping multilingual translation in large language models, emphasizing that translation quality is shaped not only by data volume but also by structural and typological relationships between languages.

多语言翻译语言距离信息损失大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。