arXiv:2606.12881cs.CLcs.LG2026-06

用直接偏好优化微调对话模型,更简单高效。

Direct Preference Optimization for Chatbot Fine-Tuning: An Empirical Study

论文配图:Direct Preference Optimization for Chatbot Fine-Tuning: An Empirical Study
图 1 · 摘自论文原文
  • 直接使用偏好数据优化模型,无需奖励模型
  • 训练效率提升,性能接近主流方法
  • 适合追求简洁高效的对话系统开发者

我们提出一种基于直接偏好优化(DPO)的大语言模型微调方法,该方法是一种强化学习技术。实验结果表明,DPO 能简化训练流程,提升计算效率,并取得具有竞争力的性能表现。通过 BLEU、ROUGE 和余弦相似度等指标评估,模型展现出有效的学习与收敛能力,但需进一步研究以解决观察到的训练不稳定性问题。

原文摘要 · Abstract (English)

We present an approach to fine-tuning large language models using Direct Preference Optimization (DPO), a reinforcement learning technique. Our experimental results demonstrate that DPO simplifies the training pipeline, improves computational efficiency, and achieves competitive performance. The evaluation using BLEU, ROUGE, and cosine similarity metrics indicates effective learning and convergence, though further investigation is needed to address observed training instability.

对话模型偏好优化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。