提出新攻击方法,绕过大模型自杀自残内容过滤
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
- 设计多步提示攻击,突破安全过滤机制
- 六款主流模型均被成功绕过,生成具体伤害性内容
- 揭示安全防护局限性,呼吁持续对抗测试
大型语言模型的安全防护虽不断强化,但仍易受新型对抗性提示攻击。本文针对自杀与自残场景,提出两种新的多步提示级越狱攻击方法,可有效绕过内置内容与安全过滤机制。实验在六款广泛使用的LLM上验证,结果表明用户意图被忽略,模型生成了详细且具有现实危害的有害内容。研究揭示了提示-响应过滤机制的多重伦理困境,并强调在安全关键型AI部署中需开展持续的对抗性测试。尽管应建立明确的安全措施,但当前通用大模型的技术成熟度下,实现全场景、全领域的全面安全仍极为困难。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have led to increasingly sophisticated safety protocols and features designed to prevent harmful, unethical, or unauthorized outputs. However, these guardrails remain susceptible to novel and creative forms of adversarial prompting, including manually generated test cases. In this work, we present two new test cases in mental health for (i) suicide and (ii) self-harm, using multi-step, prompt-level jailbreaking and bypass built-in content and safety filters. We show that user intent is disregarded, leading to the generation of detailed harmful content and instructions that could cause real-world harm. We conduct an empirical evaluation across six widely available LLMs, demonstrating the generalizability and reliability of the bypass. We assess these findings and the multilayered ethical tensions that they present for their implications on prompt-response filtering and context- and task-specific model development. We recommend a more comprehensive and systematic approach to AI safety and ethics while emphasizing the need for continuous adversarial testing in safety-critical AI deployments. We also argue that while certain clearly defined safety measures and guardrails can and must be implemented in LLMs, ensuring robust and comprehensive safety across all use cases and domains remains extremely challenging given the current technical maturity of general-purpose LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。