研究思维链模型在微调攻击下的安全漏洞
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
- 用微调攻击诱导思维链模型生成有害内容
- 攻击使模型输出危害性显著增强
- 揭示思维链模型在安全部署中的风险
大语言模型在预训练阶段可能学习到大量潜在有害信息。微调攻击可利用这些信息,诱导模型暴露不良行为,从而生成有害内容。本文聚焦于基于思维链推理的模型 DeepSeek,研究其在微调攻击下的表现。具体探讨微调如何操纵模型输出,加剧响应的危害性,并分析思维链推理与对抗性输入之间的交互作用。本研究旨在揭示思维链模型对微调攻击的脆弱性及其在安全与伦理部署中的影响。
原文摘要 · Abstract (English)
Large language models are typically trained on vast amounts of data during the pre-training phase, which may include some potentially harmful information. Fine-tuning attacks can exploit this by prompting the model to reveal such behaviours, leading to the generation of harmful content. In this paper, we focus on investigating the performance of the Chain of Thought based reasoning model, DeepSeek, when subjected to fine-tuning attacks. Specifically, we explore how fine-tuning manipulates the model's output, exacerbating the harmfulness of its responses while examining the interaction between the Chain of Thought reasoning and adversarial inputs. Through this study, we aim to shed light on the vulnerability of Chain of Thought enabled models to fine-tuning attacks and the implications for their safety and ethical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。