arXiv:2412.00074cs.CL2024-12被引 2

通过指令调优提升大模型安全与有用性,效果显著。

Safe to Serve: Aligning Instruction-Tuned Models for Safety and Helpfulness

  • 在指令调优中加入安全指令,平衡模型表现与安全性。
  • 安全回复率从40%提升至90%以上,优于多种基线方法。
  • 适合关注大模型伦理安全的开发者与研究者使用。

大型语言模型(LLMs)在复杂推理和文本生成方面展现出卓越能力,但面对不当输入时可能生成不安全或带有偏见的内容,对实际部署带来重大伦理与实践挑战。本研究致力于构建既能提供帮助又确保无害的语言模型,解决性能与安全之间的平衡难题。研究表明,在预训练模型的指令调优过程中融入安全相关指令,可显著减少对有害提示的毒性响应,且不影响对有益性数据集的表现。其中,直接偏好优化(DPO)尤为有效,通过同时利用选择和拒绝样本进行学习,优于SIT和RAFT方法。该方法使多个有害性基准上的安全回复率从40%提升至超过90%。此外,研究还提出一套严谨的评估框架,包含专用指标与多样化数据集,全面评估模型在安全性和有用性任务中的表现。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning and text generation. However, these models can inadvertently generate unsafe or biased responses when prompted with problematic inputs, raising significant ethical and practical concerns for real-world deployment. This research addresses the critical challenge of developing language models that generate both helpful and harmless content, navigating the delicate balance between model performance and safety. We demonstrate that incorporating safety-related instructions during the instruction-tuning of pre-trained models significantly reduces toxic responses to unsafe prompts without compromising performance on helpfulness datasets. We found Direct Preference Optimization (DPO) to be particularly effective, outperforming both SIT and RAFT by leveraging both chosen and rejected responses for learning. Our approach increased safe responses from 40$\%$ to over 90$\%$ across various harmfulness benchmarks. In addition, we discuss a rigorous evaluation framework encompassing specialized metrics and diverse datasets for safety and helpfulness tasks ensuring a comprehensive assessment of the model's capabilities.

大模型安全指令调优偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。