arXiv:2605.27545cs.CL2026-05

用过去时改写指令,轻松突破多模态AI安全防线

PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI

论文配图:PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI
图 1 · 摘自论文原文
  • 通过逐步强化历史语境,动态调整提示词绕过安全限制
  • 黑盒无梯度攻击下成功率最高达100%,跨模型迁移超50%
  • 揭示安全机制根本缺陷,适合安全研究者与红队测试

多模态AI系统的越狱攻击仍鲜受关注,而不当图像生成的后果可能比文本更严重,当前防御机制也相对薄弱。我们提出PAST2HARM,一种简单但高效的自适应越狱框架,可绕过当前顶尖多模态文生图模型中的拒绝训练。基于过往研究发现,过去时改写能规避安全检测,本框架系统性利用这一漏洞。从广度上,通过时间深度化逐步增强历史锚定与档案线索,削弱不同对齐强度模型的拒绝边界;从深度上,初始合规后迭代升级,探测有害生成上限,以语言模型为裁判使用标量越狱严重度指标衡量。我们发现对话中段是峰值脆弱窗口,有害性先上升后趋于平稳,最终出现语义反转。在Gemini Nano Banana Pro、GPT Image 2和SD XL三款模型上,黑盒无梯度设置下攻击成功率分别为83%、67%和100%。对抗性提示具有跨模型迁移能力,交叉成功率超50%。攻击诱发了包括露骨性内容、政治虚假信息、历史否认叙事、仇恨言论及自残美化在内的多样有害输出。我们还发布了包含提示、改写和输出的精选基准数据集,供红队测试与对齐研究使用。结果揭示当前安全机制存在根本脆弱性,亟需更强的多模态安全训练。

原文摘要 · Abstract (English)

Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text and current defenses are relatively immature. We introduce PAST2HARM, a simple yet effective adaptive jailbreak framework that bypasses refusal training in state of the art multimodal text to image models. Building on prior findings that past tense reformulations can evade safeguards, PAST2HARM systematically exploits this vulnerability in multimodal generative AI. We characterize the attack along two dimensions. First, breadth: through temporal deepening, the framework incrementally strengthens historical anchoring and archival cues, eroding refusal boundaries across models with varying alignment strength. Second, depth: via iterative escalation after initial compliance, we probe the upper bound of harmful generation, measuring severity using a scalar severity jailbreak metric evaluated by a language model acting as a judge. We find that mid conversation turns form peak vulnerability windows, where harmfulness increases before plateauing and eventually undergoing semantic inversion. We evaluate PAST2HARM on three models Gemini Nano Banana Pro, GPT Image 2, and SD XL achieving attack success rates of 83 percent, 67 percent, and 100 percent in a black box, gradient free setting. Adversarial prompts also transfer across models, with cross model success rates above 50 percent. The attack elicits diverse harmful outputs, including explicit sexual content, political disinformation, historical denial narratives, hate speech, and self harm glorification. We further release a curated benchmark of prompts, reformulations, and outputs as a resource for red teaming and alignment. Our results expose fundamental brittleness in current safeguards and highlight the need for stronger multimodal safety training.

越狱攻击多模态安全对抗样本红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。