arXiv:2511.10519cs.CLcs.AI2025-11被引 2

用不同语言风格可绕过AI安全防护,让模型说出不当内容。

Say It Differently: Linguistic Styles as Jailbreak Vectors

  • 通过11种语言风格重写提示词,保持原意但改变表达方式。
  • 某些风格使攻击成功率最高提升57个百分点。
  • 适合研究模型安全、对抗攻击或内容过滤的团队参考。

大型语言模型通常评估对改写或语义等价攻击的鲁棒性,但语言风格变化作为攻击面却少受关注。本文系统研究恐惧、好奇等语言风格如何重构有害意图,诱导对齐模型产生不安全响应。我们基于3个标准数据集,利用手工模板和大模型重写,将提示词转化为11种不同语言风格,同时保持语义一致,构建了风格增强的越狱基准。评估16个开源与闭源指令微调模型发现,风格重写使越狱成功率最高提升57个百分点。其中恐惧、好奇和共情类风格最有效,上下文相关的重写优于固定模板。为缓解此问题,我们提出使用二级大模型进行风格中性化预处理,剥离操纵性语言特征,显著降低越狱成功率。研究揭示了当前安全流程中被忽视的系统性且可扩展的漏洞。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are commonly evaluated for robustness against paraphrased or semantically equivalent jailbreak prompts, yet little attention has been paid to linguistic variation as an attack surface. In this work, we systematically study how linguistic styles such as fear or curiosity can reframe harmful intent and elicit unsafe responses from aligned models. We construct style-augmented jailbreak benchmark by transforming prompts from 3 standard datasets into 11 distinct linguistic styles using handcrafted templates and LLM-based rewrites, while preserving semantic intent. Evaluating 16 open- and close-source instruction-tuned models, we find that stylistic reframing increases jailbreak success rates by up to +57 percentage points. Styles such as fearful, curious and compassionate are most effective and contextualized rewrites outperform templated variants. To mitigate this, we introduce a style neutralization preprocessing step using a secondary LLM to strip manipulative stylistic cues from user inputs, significantly reducing jailbreak success rates. Our findings reveal a systemic and scaling-resistant vulnerability overlooked in current safety pipelines.

语言风格越狱攻击模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。