让翻译评估同时看多个译文,提升自动化评分准确性
COMET-poly: Machine Translation Metric Grounded in Other Candidates
- 引入多个候选译文对比,增强评估上下文信息
- 加入额外译文后相关性提升至0.118(Kendall's tau-b)
- 适合需要更贴近人工评分的翻译质量评估场景
机器翻译的自动化评估指标通常仅基于源句和单一译文,而人类评估常参考多个候选译文。这种差异可能影响评估效果。本文提出两种新指标:COMET-polycand 利用同一源句的其他译文进行对比,实现更全面的评价;COMET-polyic 受检索式上下文学习启发,引入相似源句及其人工评分,指导当前译文评估。实验显示,COMET-polycand 加入一个额外译文后,段级评估相关性从0.079提升至0.118,增加译文数量仍有提升;COMET-polyic 使用检索示例也取得类似改进(0.079 → 0.116)。模型已公开。
原文摘要 · Abstract (English)
Automated metrics for machine translation attempt to replicate human judgment. Unlike humans, who often assess a translation in the context of multiple alternatives, these metrics typically consider only the source sentence and a single translation. This discrepancy in the evaluation setup may negatively impact the performance of automated metrics. We propose two automated metrics that incorporate additional information beyond the single translation. COMET-polycand uses alternative translations of the same source sentence to compare and contrast with the translation at hand, thereby providing a more informed assessment of its quality. COMET-polyic, inspired by retrieval-based in-context learning, takes in translations of similar source texts along with their human-labeled quality scores to guide the evaluation. We find that including a single additional translation in COMET-polycand improves the segment-level metric performance (0.079 to 0.118 Kendall's tau-b correlation), with further gains when more translations are added. Incorporating retrieved examples in COMET-polyic yields similar improvements (0.079 to 0.116 Kendall's tau-b correlation). We release our models publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。