arXiv:2409.17673cs.CL2024-09被引 6

用人类偏好优化翻译模型,跨语言提升译文质量

Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization

  • 用质量评估模型替代人工评分,直接优化翻译质量
  • 仅对部分语言做优化,全语言性能均提升
  • 适合多语言翻译系统改进与自动评估研究者

基于人类反馈的强化学习(RLHF)及其衍生方法如直接偏好优化(DPO),是将通用基础模型适配特定任务的任务对齐算法。本文表明,将任务对齐应用于神经机器翻译(NMT)可缓解现有任务-数据不匹配问题,即使仅对部分语言进行任务对齐,多语言模型所有语言的性能均得到提升。为此,我们提出直接质量优化(DQO),一种利用预训练翻译质量评估模型作为人类偏好代理的DPO变体,并通过自动指标和人工评估验证了其有效性。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) and derivative techniques like Direct Preference Optimization (DPO) are task-alignment algorithms used to repurpose general, foundational models for specific tasks. We show that applying task-alignment to neural machine translation (NMT) addresses an existing task--data mismatch in NMT, leading to improvements across all languages of a multilingual model, even when task-alignment is only applied to a subset of those languages. We do so by introducing Direct Quality Optimization (DQO), a variant of DPO leveraging a pre-trained translation quality estimation model as a proxy for human preferences, and verify the improvements with both automatic metrics and human evaluation.

机器翻译偏好优化多语言质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。