普通人用巧妙提示词就能绕过大模型安全机制。
Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
- 用多轮叙事、伪装词汇等技巧构造提示词
- 实测发现所有安全环节都可被低门槛攻破
- 适合关注模型安全与对抗攻击的研究者
尽管大型语言模型(LLMs)和文本到图像(T2I)系统在对齐与内容审核方面取得显著进展,但仍易受提示词攻击(即“越狱”)影响。与需要专业知识的传统对抗样本不同,当前许多越狱攻击由普通用户仅通过精心设计的提示词即可实现。本文通过系统性研究,揭示了非专家如何利用多轮叙事升级、词汇伪装、暗示链、虚构身份扮演及细微语义修改等技术,可靠地绕过安全机制。我们提出一个涵盖文本输出与T2I模型的统一提示级越狱策略分类框架,并基于主流API的实证案例进行验证。分析表明,从输入过滤到输出校验的每个安全环节均可被低成本策略突破。研究强调亟需具备上下文感知能力的防御机制,以应对这些越狱在真实场景中易于复现的威胁。
原文摘要 · Abstract (English)
Despite significant advancements in alignment and content moderation, large language models (LLMs) and text-to-image (T2I) systems remain vulnerable to prompt-based attacks known as jailbreaks. Unlike traditional adversarial examples requiring expert knowledge, many of today's jailbreaks are low-effort, high-impact crafted by everyday users with nothing more than cleverly worded prompts. This paper presents a systems-style investigation into how non-experts reliably circumvent safety mechanisms through techniques such as multi-turn narrative escalation, lexical camouflage, implication chaining, fictional impersonation, and subtle semantic edits. We propose a unified taxonomy of prompt-level jailbreak strategies spanning both text-output and T2I models, grounded in empirical case studies across popular APIs. Our analysis reveals that every stage of the moderation pipeline, from input filtering to output validation, can be bypassed with accessible strategies. We conclude by highlighting the urgent need for context-aware defenses that reflect the ease with which these jailbreaks can be reproduced in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。