arXiv:2604.12911cs.CLcs.AI2026-04

用往返翻译检测模型真外语能力,比现有评测更靠谱。

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

  • 用原文转目标语言再转回的往返翻译测多语生成缺陷。
  • 往返翻译与真实任务评分相关性高达0.94,几乎完美匹配。
  • 无需人工参考译文,适合评估前沿多语模型性能。

多语言基准测试引导前沿模型发展,但现有评测方式与主流推理和知识类基准类似,仅跨语言展开。我们发现这些评测实际测量的是数学推理与事实记忆能力,而非真正的多语言能力。例如,思考型模型在这些评测中表现远超指令型,但在真实任务(如LMArena)中反而更差。为此,我们提出往返翻译评测:将源语言文本翻译到目标语言再译回,原始与结果间的语义偏差暴露模型多语生成缺陷。该方法与LMArena用户评分高度相关(r = 0.94),无需人工参考译文,也不需比被测模型更强的多语判别器。最后,我们发布跨全球广泛使用语言的挑战性基准LiT,用于真实评估前沿多语模型。

原文摘要 · Abstract (English)

Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to popular reasoning and knowledge benchmarks, but across many languages. We show such benchmarks, and consequently multilingual evaluations, measure mathematical reasoning and factual recall, not multilingual proficiency. For example, thinking variants dramatically outperform instruct variants on these benchmarks, yet often perform worse on real-world multilingual tasks, such as LMArena. We propose a simple alternative: evaluate multilingual capability via round-trip translation. Given text in a source language, translate it to a target language and back; semantic gaps between the original and result expose failures in multilingual generation capabilities. Round-trip translation correlates almost perfectly (\r{ho} = 0.94) with user ratings on LMArena with our benchmark, requires no human reference translations, and does not require a more capable multilingual judge than tested models. Lastly, we introduce Lost in Translation (LiT), a challenging round-trip translation benchmark spanning widely spoken languages worldwide, for realistic evaluation of multilingual frontier models.

多语言评测翻译基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。