arXiv:2605.07533cs.CL2026-05中稿 · the 26th Annual Co…

揭示大模型翻译失败的根源:低资源语言词元利用效率低

Why do Large Language Models Fail in Low-resource Translation? Unraveling the Token Dynamics of Large Language Models for Machine Translation

论文配图:Why do Large Language Models Fail in Low-resource Translation? Unraveling the Token Dynamics of Large Language Models for Machine Translation
图 1 · 摘自论文原文
  • 提出词元激活率(TAR)衡量模型对语言特有词元的使用效率
  • 发现非英语主导语言对的TAR值更低,翻译质量也更差
  • 推理型大模型在低TAR语言中生成更多词元,但效果因模型而异

大语言模型(LLMs)在机器翻译中表现强劲,但多数研究仅关注提升或评估翻译质量,缺乏对其失效原因的深入理解。本文系统分析了15个模型(含4个推理类大模型)在22种语言对上的翻译失败模式,涵盖不同资源水平的语言对。结果发现,非英语主导语言对的COMET得分普遍低于英语主导语言对。为探究成因,我们引入词元激活率(TAR),量化模型在生成过程中对语言特定词元的利用效率。通过具有已知训练数据语言分布的模型验证,TAR与语言表征能力高度相关,且较低的TAR显著关联较差的翻译性能。此外,推理类模型在翻译至低TAR语言时倾向于生成更多词元,暗示存在补偿机制,但其对翻译质量的影响因模型而异。整体表明,词元层面的动力学特征对理解大模型翻译性能至关重要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have recently demonstrated strong performance in machine translation (MT). However, most prior work focuses on improving or benchmarking translation quality, offering limited insight into when and why LLM-based translation fails. In this work, we systematically analyze failure modes of LLMs in MT by evaluating 15 models, including four reasoning LLMs, across 22 language pairs (LPs) with varying resource levels. We find that non-English-centric LPs consistently yield lower COMET scores than English-centric pairs. To investigate the underlying causes, we introduce Token Activation Rate (TAR), a metric that captures how effectively a model utilizes language-specific tokens in its vocabulary during generation. We validate TAR as a proxy for language representation using models with known language distributions in the training data, and show that lower TAR is strongly associated with poorer translation performance. Furthermore, reasoning LLMs tend to generate more tokens when translating into low-TAR languages, suggesting a compensatory mechanism, although its impact on translation quality varies across models. Overall, our findings emphasize the importance of token-level dynamics in understanding MT performance of LLMs.

大模型翻译词元动态低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。