arXiv:2511.21718cs.CL2025-11

用隐含概念触发让大模型输出有害内容,突破安全防护

When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers

  • 通过特定短语构造概念触发,操纵模型隐含价值倾向
  • 在5个主流大模型上成功率高,且极少被安全机制拦截
  • 揭示商业模型对价值观层面攻击的脆弱性,适合安全研究者关注

当前大语言模型(LLM)越狱研究多聚焦于直接诱导明显有害输出,却忽视了利用模型抽象泛化能力进行的隐蔽攻击。本文提出MICM——一种无需依赖模型结构的越狱方法,针对的是模型输出中体现的聚合价值体系。基于概念形态学理论,MICM将一系列微妙概念配置编码至固定提示模板中,通过预定义短语作为概念触发器,引导模型输出特定价值立场,而不触发传统安全过滤。我们在GPT-4o、Deepseek-R1和Qwen3-8B等五个先进大模型上评估,结果表明MICM显著优于现有越狱技术,在极低拒绝率下保持高成功率。研究揭示了商用大模型在价值对齐层面存在严重漏洞:其安全机制仍易受深层价值观操控。

原文摘要 · Abstract (English)

Recent research on large language model (LLM) jailbreaks has primarily focused on techniques that bypass safety mechanisms to elicit overtly harmful outputs. However, such efforts often overlook attacks that exploit the model's capacity for abstract generalization, creating a critical blind spot in current alignment strategies. This gap enables adversaries to induce objectionable content by subtly manipulating the implicit social values embedded in model outputs. In this paper, we introduce MICM, a novel, model-agnostic jailbreak method that targets the aggregate value structure reflected in LLM responses. Drawing on conceptual morphology theory, MICM encodes specific configurations of nuanced concepts into a fixed prompt template through a predefined set of phrases. These phrases act as conceptual triggers, steering model outputs toward a specific value stance without triggering conventional safety filters. We evaluate MICM across five advanced LLMs, including GPT-4o, Deepseek-R1, and Qwen3-8B. Experimental results show that MICM consistently outperforms state-of-the-art jailbreak techniques, achieving high success rates with minimal rejection. Our findings reveal a critical vulnerability in commercial LLMs: their safety mechanisms remain susceptible to covert manipulation of underlying value alignment.

大模型安全越狱攻击价值观对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。