arXiv:2608.09551cs.CL2026-08

利用语言隐含语用背景,突破大模型安全防护

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

论文配图:Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models
图 1 · 摘自论文原文
  • 通过挖掘人类语言中未明说的隐含语境发起攻击
  • 在多个开源和闭源模型上实现高成功率攻击
  • 揭示现有对齐算法对隐含上下文的忽视问题

在大语言模型时代,攻击者常通过操控自然语言诱导模型输出有害内容,形成独特的自然语言攻击面。此类攻击依赖显式提示词,易被现有安全对齐算法防御。然而,人类语言本质上依赖语用学,需依赖世界知识、社会规范等隐含上下文理解语义。这些上下文通常未直接表达,且未被充分纳入安全对齐机制,造成人脑理解与模型安全机制的根本错位。本文首次揭示此错位带来的漏洞,定义为‘语用攻击面’。实验表明,该方法在多种开源与闭源模型上显著优于基线攻击手段,展现出极高攻击成功率。

原文摘要 · Abstract (English)

In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks directly exploit explicit linguistic cues in user prompts to bypass the safety mechanism of LLMs. However, such attacks can often be mitigated by existing safety alignment algorithms. On the other hand, human language is inherently grounded in pragmatics, necessitating typical context to interpret language, e.g., world knowledge, social norms. However, such contexts are often implicit because they are not directly expressed in human language and are not sufficiently leveraged in safety alignment, creating a fundamental mismatch between human language interpretation and safety alignment approaches. In this paper, we demonstrate that this mismatch exposes vulnerabilities in LLMs. We refer to this vulnerability as the pragmatic attack surface, which can be exploited to achieve high attack success rates. The experimental results demonstrate that our proposed approach outperforms baseline attack methods across various open-source and closed-source models by a substantial margin.

大模型安全语用漏洞攻击面

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。