arXiv:2501.17030cs.LGcs.AI2025-01被引 20

DeepSeek-R1的AI安全问题暴露了强化学习的局限性,混合训练可提升安全性。

Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies

  • 用强化学习与监督微调结合,缓解有害输出。
  • 强化学习易出现奖励劫持和语言混杂问题。
  • 适合关注大模型对齐与安全部署的研究者。

大型语言模型(LLMs)在推理、对齐和特定任务表现上取得了显著进展,但确保其无害性仍是关键挑战,尤其在深度求索R1(DeepSeek-R1)等先进模型中尤为突出。本文分析了以强化学习(RL)为主要手段减少有害输出的局限性,并与监督微调(SFT)进行对比。尽管强化学习提升了推理能力,但仍面临奖励劫持、泛化失败、语言混杂及高计算成本等问题。为此,我们提出结合强化学习与监督微调的混合训练策略,以实现更稳健的无害性降低。同时,论文还给出了使用建议及负责任部署DeepSeek-R1的未来方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable progress in reasoning, alignment, and task-specific performance. However, ensuring harmlessness in these systems remains a critical challenge, particularly in advanced models like DeepSeek-R1. This paper examines the limitations of Reinforcement Learning (RL) as the primary approach for reducing harmful outputs in DeepSeek-R1 and compares it with Supervised Fine-Tuning (SFT). While RL improves reasoning capabilities, it faces challenges such as reward hacking, generalization failures, language mixing, and high computational costs. We propose hybrid training approaches combining RL and SFT to achieve robust harmlessness reduction. Usage recommendations and future directions for deploying DeepSeek-R1 responsibly are also presented.

大模型安全强化学习对齐技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。