arXiv:2501.13927cs.CLcs.AI2025-01ACL被引 4

用置信度和奖励联合筛选难例,提升机器翻译的训练效率与质量。

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

  • 结合模型置信度与奖励分数筛选难句对,优化微调数据选择
  • 在翻译准确率和数据效率上优于RS-DPO、RSO等现有方法
  • 适用于大模型与编码器-解码器结构,如NLLB,通用性强

大语言模型在自然语言处理中展现出巨大潜力,但其在机器翻译中的应用仍面临挑战,主要源于预训练数据以英语为中心以及人类反馈强化学习(RLHF)的复杂性。直接偏好优化(DPO)作为更简单高效的替代方案,其性能高度依赖偏好数据质量。为此,本文提出置信度-奖励驱动偏好优化(CRPO),通过融合奖励得分与模型置信度,改进微调数据的选择。CRPO聚焦于模型不确定或表现不佳的困难句对,实现更有效的学习。尽管主要面向大语言模型设计,该方法也适用于NLLB等编码器-解码器架构,展现出良好泛化能力。实验表明,CRPO在翻译准确率和数据效率上均优于RS-DPO、RSO和MBR评分等现有方法。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Optimization (DPO) has emerged as a simpler and more efficient alternative, but its performance depends heavily on the quality of preference data. To address this, we propose Confidence-Reward driven Preference Optimization (CRPO), a novel method that combines reward scores with model confidence to improve data selection for fine-tuning. CRPO selects challenging sentence pairs where the model is uncertain or underperforms, leading to more effective learning. While primarily designed for LLMs, CRPO also generalizes to encoder-decoder models like NLLB, demonstrating its versatility. Empirical results show that CRPO outperforms existing methods such as RS-DPO, RSO and MBR score in both translation accuracy and data efficiency.

机器翻译偏好优化数据筛选大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。