arXiv:2609.02215cs.AI2026-09

用ASCII艺术伪装恶意请求,突破大模型安全防线

ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models

论文配图:ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
图 1 · 摘自论文原文
  • 将恶意请求嵌入可读的ASCII艺术中,伪装成艺术作品提问
  • 在11个模型上,62%的伪装请求被判定为有害,最高成功率93%
  • 攻击对模型敏感度高于话题,且规模越大越有效

安全对齐训练使大模型拒绝表面明显的有害请求,但对重新语境化的内容覆盖不足。本文提出ASCII Attack:单轮黑盒攻击,将完整可读的有害请求以ASCII艺术形式呈现为艺术品,并请求反馈。与ArtPrompt不同,该方法不隐藏内容,请求始终保持可读。模型回复以艺术评论形式出现,可能包含本应被拒绝的操作细节。每个伪装提示均配有一个直接提问的对照组,确保对比仅受语境影响。在11个模型、8类危害主题下,危害感知分类器判断62%的伪装提示为有害,而对照组为42%。最脆弱模型中,伪装提示成功率达93%。单一查询在五种危害判断标准下,有四种超过已有单次查询攻击表现。效果更依赖模型而非主题,且随模型规模增强而不减弱。近三分之二的伪装样本中,至少一名评委与多数意见相左,反映评估一致性问题,暗示模型存在泛化不匹配现象。

原文摘要 · Abstract (English)

Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefore only weakly covered. The ASCII Attack is one such recontextualisation. It is single-turn and black-box: one message, with no access to model internals. It embeds a fully legible harmful request in ASCIl-art characters, presents it as artwork, and asks for feedback. Unlike ArtPrompt, it hides nothing: the request stays readable. The reply is written as artistic critique and can contain operational detail that a plain request would have been refused for. Every framed prompt is paired with a direct-question control, so the contrast is isolated from topic, model and decoding variation. The contrast identifies a bundled surface, not one isolated channel. Across eleven models and eight harm topics, a harm-aware classifier judges 62% of framed prompts harmful against 42% of controls. On the most susceptible model the framed prompt succeeds 93% of the time. A single query matches or exceeds published single-query attacks under four of five harm judges. The effect tracks the model more than the topic and does not diminish with scale. At least one judge dissents from the panel majority on nearly two-thirds of framed rows, which is itself a measurement-validity finding. That pattern is consistent with mismatched generalisation.

安全攻防语言模型对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。