arXiv:2604.14934cs.CL2026-04ACL被引 1

构建跨语言平行质量数据集,揭示翻译评估指标的偏见问题

XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics

论文配图:XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics
图 1 · 摘自论文原文
  • 通过自动注入错误并人工筛选,生成可控质量的伪译文
  • 实证发现现有评估指标在不同语言间评分不一致
  • 提出归一化策略提升多语言评估公平性,适合评测研究者

自动评估指标对构建多语言翻译系统至关重要。当前普遍做法是跨语言平均指标得分,但这种方法存疑,因评估指标可能受跨语言评分偏差影响——相同质量的译文在不同语言中得分不同。这一问题未被系统研究,因缺乏提供跨语言平行质量实例的基准数据集,且专家标注不可行。本文提出XQ-MEval,一个半自动生成的九方向翻译数据集,用于评估翻译指标。具体方法为:自动向原文本注入MQM定义的错误,经母语者筛选确保可靠性,并合并错误生成可控质量的伪译文。这些伪译文与源句和参考译文构成三元组,用于评估指标表现。基于XQ-MEval,我们对九种代表性指标的实验揭示了平均评分与人工判断之间的不一致,首次提供了跨语言评分偏差的实证证据。最后,我们提出一种基于XQ-MEval的归一化策略,使各语言评分分布对齐,显著提升多语言评估的公平性与可靠性。

原文摘要 · Abstract (English)

Automatic evaluation metrics are essential for building multilingual translation systems. The common practice of evaluating these systems is averaging metric scores across languages, yet this is suspicious since metrics may suffer from cross-lingual scoring bias, where translations of equal quality receive different scores across languages. This problem has not been systematically studied because no benchmark exists that provides parallel-quality instances across languages, and expert annotation is not realistic. In this work, we propose XQ-MEval, a semi-automatically built dataset covering nine translation directions, to benchmark translation metrics. Specifically, we inject MQM-defined errors into gold translations automatically, filter them by native speakers for reliability, and merge errors to generate pseudo translations with controllable quality. These pseudo translations are then paired with corresponding sources and references to form triplets used in assessing the qualities of translation metrics. Using XQ-MEval, our experiments on nine representative metrics reveal the inconsistency between averaging and human judgment and provide the first empirical evidence of cross-lingual scoring bias. Finally, we propose a normalization strategy derived from XQ-MEval that aligns score distributions across languages, improving the fairness and reliability of multilingual metric evaluation.

机器翻译评估指标跨语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。