arXiv:2411.00390cs.CLcs.AI2024-11被引 9

用人类偏好校准模型,让翻译评估更贴近真实判断。

MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration

  • 基于高斯过程的贝叶斯优化,动态调整评估指标
  • 在WMT24数据集上超越所有现有基准
  • 兼顾效率与准确性,适合参考句和无参考句场景

我们提出MetaMetrics-MT,一种通过贝叶斯优化结合高斯过程与人类偏好校准的机器翻译评估新方法。该方法优化现有评估指标与人工判断的相关性,在WMT24基准任务数据集上的实验表明,其在有参考句设置下优于所有现有基线,达到当前最佳性能;在无参考句设置下表现媲美领先指标,同时具备更高效率。

原文摘要 · Abstract (English)

We present MetaMetrics-MT, an innovative metric designed to evaluate machine translation (MT) tasks by aligning closely with human preferences through Bayesian optimization with Gaussian Processes. MetaMetrics-MT enhances existing MT metrics by optimizing their correlation with human judgments. Our experiments on the WMT24 metric shared task dataset demonstrate that MetaMetrics-MT outperforms all existing baselines, setting a new benchmark for state-of-the-art performance in the reference-based setting. Furthermore, it achieves comparable results to leading metrics in the reference-free setting, offering greater efficiency.

机器翻译评估指标人类偏好贝叶斯优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。