arXiv:2508.14909cs.CL2025-08被引 3

WMT25机器翻译初评结果出炉,助力参赛者撰写系统论文。

Preliminary Ranking of WMT25 General Machine Translation Systems

  • 基于自动评估指标对参赛系统进行初步排序
  • 排名可能偏向使用重排序技术的系统
  • 结果仅作参考,最终以人工评测为准

我们公布了提交至WMT25通用机器翻译共享任务的系统初步排名,该排名由自动评估指标得出。由于依赖自动评价,结果可能偏向采用重排序技术(如质量估计或最小贝叶斯风险解码)的系统。官方最终排名将基于人工评估,更具可靠性,并将取代此初步结果。提前发布这些发现旨在协助参赛者撰写系统描述论文,而非提供最终结论。

原文摘要 · Abstract (English)

We present the preliminary rankings of machine translation (MT) systems submitted to the WMT25 General Machine Translation Shared Task, as determined by automatic evaluation metrics. Because these rankings are derived from automatic evaluation, they may exhibit a bias toward systems that employ re-ranking techniques, such as Quality Estimation or Minimum Bayes Risk decoding. The official WMT25 ranking will be based on human evaluation, which is more reliable and will supersede these results. The official WMT25 ranking will be based on human evaluation, which is more reliable and will supersede these results. The purpose of releasing these findings now is to assist task participants with their system description papers; not to provide final findings.

机器翻译自动评估WMT25

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。