用多个评估指标组合解码,可减少因单一指标导致的翻译质量误判。
Mitigating Metric Bias in Minimum Bayes Risk Decoding
- 采用多个神经评估指标的集成作为解码目标,避免单一指标偏差。
- 实验显示,集成指标解码在人工评估中优于单指标方法。
- 该方法适合追求高精度机器翻译质量评估的研究者和开发者。
使用如COMET或MetricX等指标进行最小贝叶斯风险(MBR)解码,虽优于贪心或束搜索等传统方法,但会引入所谓‘指标偏差’问题。由于MBR解码旨在最大化特定指标得分,导致无法使用同一指标同时进行解码与评估,因为得分提升可能源于奖励黑客而非真实质量改进。本文发现,相比人类评分,神经指标在使用相同指标作为效用函数时,不仅高估了MBR解码的质量,也高估了使用其他神经效用指标的MBR/QE解码质量。我们进一步证明,通过在MBR解码中使用效用指标的集成,可有效缓解指标偏差:人工评估显示,基于集成指标的解码性能优于单一指标。
原文摘要 · Abstract (English)
While Minimum Bayes Risk (MBR) decoding using metrics such as COMET or MetricX has outperformed traditional decoding methods such as greedy or beam search, it introduces a challenge we refer to as metric bias. As MBR decoding aims to produce translations that score highly according to a specific utility metric, this very process makes it impossible to use the same metric for both decoding and evaluation, as improvements might simply be due to reward hacking rather than reflecting real quality improvements. In this work we find that compared to human ratings, neural metrics not only overestimate the quality of MBR decoding when the same metric is used as the utility metric, but they also overestimate the quality of MBR/QE decoding with other neural utility metrics as well. We also show that the metric bias issue can be mitigated by using an ensemble of utility metrics during MBR decoding: human evaluations show that MBR decoding using an ensemble of utility metrics outperforms a single utility metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。