arXiv:2410.07779cs.CL2024-10EMNLP被引 12

用自动评估指标构建翻译偏好数据集,提升机器翻译质量

Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation

  • 结合专业译员评分与自动指标,构建混合偏好数据
  • 新数据集含1.8万条跨18语种翻译样本,覆盖2022年后多领域文本
  • 在WMT23和FLORES测试中显著提升TOWER模型表现

对齐人类偏好是训练准确且安全大语言模型的关键步骤,机器翻译(MT)同样如此。更精准处理语言细节和上下文变化可提升翻译质量。然而,基于人工反馈的偏好数据获取与整理成本高昂。自动评估指标虽可生成偏好信号,但未必完全符合人类判断。本文提出一种融合二者优势的方法:首先由专业语言学家对多个高质量翻译系统输出的句子进行质量评分,并评估现有自动指标在恢复这些偏好方面的表现;随后基于分析结果,构建新的数据集MT-Pref(基于指标诱导的翻译偏好数据集),包含18,000个实例,覆盖18种语言方向,文本来源为2022年后多个领域。实验表明,在MT-Pref上对TOWER模型进行对齐训练,能显著提升其在WMT23和FLORES基准上的翻译质量。

原文摘要 · Abstract (English)

Alignment with human preferences is an important step in developing accurate and safe large language models. This is no exception in machine translation (MT), where better handling of language nuances and context-specific variations leads to improved quality. However, preference data based on human feedback can be very expensive to obtain and curate at a large scale. Automatic metrics, on the other hand, can induce preferences, but they might not match human expectations perfectly. In this paper, we propose an approach that leverages the best of both worlds. We first collect sentence-level quality assessments from professional linguists on translations generated by multiple high-quality MT systems and evaluate the ability of current automatic metrics to recover these preferences. We then use this analysis to curate a new dataset, MT-Pref (metric induced translation preference) dataset, which comprises 18k instances covering 18 language directions, using texts sourced from multiple domains post-2022. We show that aligning TOWER models on MT-Pref significantly improves translation quality on WMT23 and FLORES benchmarks.

机器翻译偏好学习自动评估数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。