arXiv:2609.04173cs.CL2026-09

打造可持续更新的翻译测试集,揭示主流模型真实缺陷。

Last Translation Benchmark

  • 用人工创作且同行评审的多模态案例挑战顶尖翻译模型
  • 每例附手工验证规则,精准定位模型失败场景
  • 支持持续贡献,适合研究模型边界与评测方法者

为推动科学进步,需要能测试前沿模型极限并揭示失败模式的基准与评估方法。随着模型能力提升,现有机器翻译基准趋于饱和,自动评估指标不可靠、易被奖励劫持,且难以提供可操作反馈。人工评估也存在可复现性差、客观性不足和扩展性差的问题,阻碍了领域进展的客观追踪与改进路径识别。本文提出「最后翻译基准」(Last Translation Benchmark),包含经人工撰写与同行评审的文本、图像、音频、视频等多模态示例,专门设计用于突破当前领先翻译模型。我们引入新评估方式:每个示例配有手工编写的确切验证规则,明确指出模型在该例中的具体失败模式,实现可靠且可行动的评估。该基准为动态数据集,持续接受新贡献。最新版本为 LTBv1,收录截至2026年9月1日前的通过内容,未来将定期发布新版本。

原文摘要 · Abstract (English)

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.

机器翻译评估基准多模态可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。