模型在道德困境中易被诱导执行有害行为,新方法可识别并防御此类攻击。
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs
- 通过嵌入道德框架的多轮对抗测试,暴露模型伦理推理漏洞。
- 攻击成功率高,因模型误将有害行为视为道德必要妥协。
- 提出新防御框架,区分支持危害与分析道德的响应,保持模型有用性。
大型语言模型的安全对齐主要基于请求安全或不安全的二元假设,但在伦理困境中此分类不足。我们提出TRIAL方法,通过多轮红队测试,将有害请求嵌入道德语境,系统性利用模型的伦理推理能力,使有害行为被视作道德必要妥协,从而实现高成功率攻击。基于此,我们提出ERR(伦理推理鲁棒性)防御框架,区分促进危害的工具性回应与不支持危害的解释性回应。该框架采用分层有害门控LoRA结构,在抵御基于推理的攻击的同时保持模型实用性。
原文摘要 · Abstract (English)
Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classification proves insufficient when models encounter ethical dilemmas, where the capacity to reason through moral trade-offs creates a distinct attack surface. We formalize this vulnerability through TRIAL, a multi-turn red-teaming methodology that embeds harmful requests within ethical framings. TRIAL achieves high attack success rates across most tested models by systematically exploiting the model's ethical reasoning capabilities to frame harmful actions as morally necessary compromises. Building on these insights, we introduce ERR (Ethical Reasoning Robustness), a defense framework that distinguishes between instrumental responses that enable harmful outcomes and explanatory responses that analyze ethical frameworks without endorsing harmful acts. ERR employs a Layer-Stratified Harm-Gated LoRA architecture, achieving robust defense against reasoning-based attacks while preserving model utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。