谷歌改进翻译评估模型,支持有无参考文本都可评分。
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
- 混合使用参考文本与无参考的评分方式,适应多种评估场景。
- 在WMT23数据和自建挑战集上,相比前版提升显著。
- 通过合成数据增强,有效应对流畅但无关或翻译不足问题。
本文介绍谷歌提交至WMT2024翻译评估共享任务的MetricX-24方案。主要提交版本为一种混合参考依赖与参考无关的评估指标,无论是否提供源句或参考译文,均可对翻译质量进行评分。该模型采用两阶段训练:第一阶段仅使用DA评分,第二阶段结合MQM与DA评分。两个阶段的训练数据均通过自动生成的合成样例进行增强,以提升对常见失败模式(如流畅但无关、翻译不全)的鲁棒性。通过消融实验验证各项改进效果,在WMT23的MQM评分及新构建的合成挑战集上均表现出显著优于MetricX-23的性能。
原文摘要 · Abstract (English)
In this paper, we present the MetricX-24 submissions to the WMT24 Metrics Shared Task and provide details on the improvements we made over the previous version of MetricX. Our primary submission is a hybrid reference-based/-free metric, which can score a translation irrespective of whether it is given the source segment, the reference, or both. The metric is trained on previous WMT data in a two-stage fashion, first on the DA ratings only, then on a mixture of MQM and DA ratings. The training set in both stages is augmented with synthetic examples that we created to make the metric more robust to several common failure modes, such as fluent but unrelated translation, or undertranslation. We demonstrate the benefits of the individual modifications via an ablation study, and show a significant performance increase over MetricX-23 on the WMT23 MQM ratings, as well as our new synthetic challenge set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。