arXiv:2512.18906cs.CL2025-12ACL被引 1

无需错误标注,生成式评估模型可解释机器翻译质量

Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations

  • 用强化学习从翻译偏好中训练,生成分步推理过程
  • 仅用6万对数据即达顶尖评分器水平,且在分布外测试中稳健
  • 能自动生成改译建议,适合需要可解释性与质量提升的场景

多年来,自动机器翻译评估指标在基准测试中不断进步,有时甚至达到人类水平。然而,它们仍是黑箱,难以解释决策过程,且在真实世界的分布外输入下常表现不佳。我们提出Remedy-R,一种基于强化学习、无需错误标注或闭源大模型蒸馏的生成式机器翻译评估方法。Remedy-R通过逐步分析准确性、流畅性和完整性生成最终得分,实现更可解释的评估。仅使用跨两种语言对的6万组训练样本,其在WMT22-24元评估中表现媲美顶级标量指标和GPT-4判官,具备良好跨语言泛化能力,并在分布外压力测试中表现出强鲁棒性。此外,其生成的自省反馈可复用于改进翻译。基于此,我们构建Remedy-R Agent——一个评估-修正流水线,利用其分析结果优化多种模型(Qwen2.5、ALMA-R、GPT-4o-mini、Gemini-2.0-Flash)的翻译质量,表明其推理机制捕捉到关键翻译信息,具有实际应用价值。

原文摘要 · Abstract (English)

Over the years, automatic MT metrics have hillclimbed benchmarks and presented strong and sometimes human-level agreement with human ratings. Yet they remain black-box, offering little insight into their decision-making and often failing under real-world out-of-distribution (OOD) inputs. We introduce Remedy-R, a reasoning-driven generative MT metric trained with reinforcement learning from pairwise translation preferences, without requiring error-span annotations or distillation from closed LLMs. Remedy-R produces step-by-step analyses of accuracy, fluency, and completeness, followed by a final score, enabling more interpretable assessments. With only 60K training pairs across two language pairs, Remedy-R remains competitive with top scalar metrics and GPT-4-based judges on WMT22-24 meta-evaluation, generalizes to other languages, and exhibits strong robustness on OOD stress tests. Moreover, Remedy-R models generate self-reflective feedback that can be reused for translation improvement. Building on this finding, we introduce Remedy-R Agent, a simple evaluate-revise pipeline that leverages Remedy-R's evaluation analysis to refine translations. This agent consistently improves translation quality across diverse models, including Qwen2.5, ALMA-R, GPT-4o-mini, and Gemini-2.0-Flash, suggesting that Remedy-R's reasoning captures translation-relevant information and is practically useful.

机器翻译生成评估可解释性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。