用人类偏好优化翻译模型,跨语言提升译文质量
Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization
- 用质量评估模型替代人工评分,直接优化翻译质量
- 仅对部分语言做优化,全语言性能均提升
- 适合多语言翻译系统改进与自动评估研究者
基于人类反馈的强化学习(RLHF)及其衍生方法如直接偏好优化(DPO),是将通用基础模型适配特定任务的任务对齐算法。本文表明,将任务对齐应用于神经机器翻译(NMT)可缓解现有任务-数据不匹配问题,即使仅对部分语言进行任务对齐,多语言模型所有语言的性能均得到提升。为此,我们提出直接质量优化(DQO),一种利用预训练翻译质量评估模型作为人类偏好代理的DPO变体,并通过自动指标和人工评估验证了其有效性。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) and derivative techniques like Direct Preference Optimization (DPO) are task-alignment algorithms used to repurpose general, foundational models for specific tasks. We show that applying task-alignment to neural machine translation (NMT) addresses an existing task--data mismatch in NMT, leading to improvements across all languages of a multilingual model, even when task-alignment is only applied to a subset of those languages. We do so by introducing Direct Quality Optimization (DQO), a variant of DPO leveraging a pre-trained translation quality estimation model as a proxy for human preferences, and verify the improvements with both automatic metrics and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。