arXiv:2509.09055cs.CLcs.AI2025-09被引 4

SFT与DPO结合提升小模型安全与助人能力

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M

  • 用SFT和DPO联合微调,互补提升模型对齐效果
  • 联合模型在安全性、帮助性上均优于单一方法
  • 适合关注模型对齐与微调策略的研究者

本研究考察了监督微调(SFT)、直接偏好优化(DPO)及两者结合(SFT+DPO)在提升OPT-350M语言模型安全性与帮助性方面的效果。基于Anthropic Helpful-Harmless RLHF数据集,训练并评估了四个模型:基础OPT350M、SFT模型、DPO模型及SFT+DPO联合模型。引入三项评价指标:无害率(HmR)、帮助性率(HpR)与综合对齐得分(CAS),均来自奖励模型输出。结果表明,尽管SFT表现优于DPO,但联合模型在所有指标上均超越其他模型,证明两种技术具有互补性。研究还揭示了噪声数据、有限GPU资源与训练约束带来的挑战。本工作为理解微调策略如何影响模型对齐提供了全面视角,并为未来更稳健的对齐流程奠定基础。

原文摘要 · Abstract (English)

This research investigates the effectiveness of alignment techniques, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and a combined SFT+DPO approach on improving the safety and helpfulness of the OPT-350M language model. Utilizing the Anthropic Helpful-Harmless RLHF dataset, we train and evaluate four models: the base OPT350M, an SFT model, a DPO model, and a model trained with both SFT and DPO. We introduce three key evaluation metrics: Harmlessness Rate (HmR), Helpfulness Rate (HpR), and a Combined Alignment Score (CAS), all derived from reward model outputs. The results show that while SFT outperforms DPO, The combined SFT+DPO model outperforms all others across all metrics, demonstrating the complementary nature of these techniques. Our findings also highlight challenges posed by noisy data, limited GPU resources, and training constraints. This study offers a comprehensive view of how fine-tuning strategies affect model alignment and provides a foundation for more robust alignment pipelines in future work.

大模型对齐SFTDPO模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。