冲突让大模型易受攻击,安全机制会失效。
Conflicts Make Large Reasoning Models Vulnerable to Attacks
- 在对立目标下测试大模型决策行为,发现安全机制被干扰。
- 冲突使攻击成功率显著上升,单轮提问即可触发漏洞。
- 适合关注大模型安全与对齐的研究者阅读。
大型推理模型在多个领域表现卓越,但其在面对冲突目标时的决策机制仍不清晰。本文研究了模型在两类冲突下的响应:内部冲突(对齐价值间的对立)和困境(相互矛盾的选择,包括牺牲型、胁迫型、代理中心型和社会型)。通过五个基准上的1300多个提示,评估了Llama-3.1-Nemotron-8B、QwQ-32B和DeepSeek R1三种代表性模型,发现冲突显著提高了攻击成功率,即使在无复杂自动化攻击的单轮非叙事提问中亦然。层间与神经元级分析显示,安全相关表征与功能表征在冲突下发生转移与重叠,干扰了安全对齐行为。本研究强调需发展更深层的对齐策略以保障下一代推理模型的鲁棒性与可信度。代码已公开于https://github.com/DataArcTech/ConflictHarm。警告:本文包含不当、冒犯及有害内容。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have achieved remarkable performance across diverse domains, yet their decision-making under conflicting objectives remains insufficiently understood. This work investigates how LRMs respond to harmful queries when confronted with two categories of conflicts: internal conflicts that pit alignment values against each other and dilemmas, which impose mutually contradictory choices, including sacrificial, duress, agent-centered, and social forms. Using over 1,300 prompts across five benchmarks, we evaluate three representative LRMs - Llama-3.1-Nemotron-8B, QwQ-32B, and DeepSeek R1 - and find that conflicts significantly increase attack success rates, even under single-round non-narrative queries without sophisticated auto-attack techniques. Our findings reveal through layerwise and neuron-level analyses that safety-related and functional representations shift and overlap under conflict, interfering with safety-aligned behavior. This study highlights the need for deeper alignment strategies to ensure the robustness and trustworthiness of next-generation reasoning models. Our code is available at https://github.com/DataArcTech/ConflictHarm. Warning: This paper contains inappropriate, offensive and harmful content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。