调长输出可提升模型安全,但需动态控制避免被攻击利用。
Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
- 通过自纠正机制增强长输出的安全性
- 部分攻击能利用生成长度延长来绕过防御
- 适合关注大模型安全与推理优化的研究者
大型语言模型(LLMs)展现出强大的推理能力,但在对抗性条件下的安全性仍存挑战。本研究考察输出长度对DeepSeek-R1在强制思考(Forced Thinking)场景下鲁棒性的影响。我们分析了多种对抗性提示下的响应,发现较长输出可通过自我校正提升安全性,但某些攻击类型会利用生成长度的延长进行突破。结果表明,应动态调控输出长度以平衡推理效能与安全。为此,我们提出基于强化学习的策略调整和自适应标记长度调节方法,以增强LLM的安全性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated strong reasoning capabilities, but their safety under adversarial conditions remains a challenge. This study examines the impact of output length on the robustness of DeepSeek-R1, particularly in Forced Thinking scenarios. We analyze responses across various adversarial prompts and find that while longer outputs can improve safety through self-correction, certain attack types exploit extended generations. Our findings suggest that output length should be dynamically controlled to balance reasoning effectiveness and security. We propose reinforcement learning-based policy adjustments and adaptive token length regulation to enhance LLM safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。