arXiv:2512.21110cs.AIcs.CL2025-12被引 2

大模型常误解用户意图,导致安全机制被绕过。

Beyond Context: Large Language Models' Failure to Grasp Users' Intent

  • 用情绪化、渐进式、学术化方式诱导模型输出
  • 开启推理模式反而让攻击更有效,准确率提升但不识别恶意意图
  • 仅Claude Opus 4.1在部分场景识别意图,凸显设计缺陷

当前大语言模型的安全机制主要针对显性有害内容,却忽视了理解上下文与识别用户意图的能力短板。这导致恶意用户可系统性绕过安全防护。我们对ChatGPT、Claude、Gemini、DeepSeek等主流LLM进行实证评估,发现通过情感包装、逐步披露和学术化理由等手段,可有效规避可靠的安全机制。值得注意的是,启用推理功能反而增强了攻击效果,虽提升了事实准确性,却未能识别潜在恶意意图。唯一例外是Claude Opus 4.1,在某些场景下优先判断意图而非提供信息。这一模式表明,现有架构存在系统性漏洞,亟需将上下文理解与意图识别作为核心安全能力,而非事后补救措施。

原文摘要 · Abstract (English)

Current Large Language Models (LLMs) safety approaches focus on explicitly harmful content while overlooking a critical vulnerability: the inability to understand context and recognize user intent. This creates exploitable vulnerabilities that malicious users can systematically leverage to circumvent safety mechanisms. We empirically evaluate multiple state-of-the-art LLMs, including ChatGPT, Claude, Gemini, and DeepSeek. Our analysis demonstrates the circumvention of reliable safety mechanisms through emotional framing, progressive revelation, and academic justification techniques. Notably, reasoning-enabled configurations amplified rather than mitigated the effectiveness of exploitation, increasing factual precision while failing to interrogate the underlying intent. The exception was Claude Opus 4.1, which prioritized intent detection over information provision in some use cases. This pattern reveals that current architectural designs create systematic vulnerabilities. These limitations require paradigmatic shifts toward contextual understanding and intent recognition as core safety capabilities rather than post-hoc protective mechanisms.

大模型安全意图识别攻击防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。