arXiv:2503.03502cs.CLcs.AI2025-03

通过几何特性检测恶意提示,高效且通用。

Geometry-Guided Adversarial Prompt Detection via Curvature and Local Intrinsic Dimension

  • 基于文本嵌入空间的曲率与局部内在维数分析
  • 对齐提示几何特征,实现近乎完美的检测准确率
  • 不依赖模型架构,适合多种大模型和攻击类型

对抗性提示能够突破前沿大语言模型(LLMs)的限制并诱导不良行为,严重阻碍其安全部署。现有缓解策略多依赖激活内置防御机制或微调模型,成本高且损害模型性能。相比之下,检测方法更高效实用。但对抗性与正常提示的根本差异尚不明确。本文提出CurvaLID框架,利用提示的几何特性实现高效检测。该方法对模型类型无感,适用于不同攻击类型和模型架构。我们基于Whewell方程将曲率概念推广至n维词嵌入空间,量化语义偏移与流形局部曲率。同时引入局部内在维数(LID)捕捉对抗子空间中的互补几何特征。实验表明,对抗性提示具有显著不同的几何特征,使CurvaLID在多个数据集上实现接近完美的分类性能,优于当前最优检测器。该方法为抵御恶意查询提供了可靠、高效的模型无关防护方案。

原文摘要 · Abstract (English)

Adversarial prompts are capable of jailbreaking frontier large language models (LLMs) and inducing undesirable behaviours, posing a significant obstacle to their safe deployment. Current mitigation strategies primarily rely on activating built-in defence mechanisms or fine-tuning LLMs, both of which are computationally expensive and can sacrifice model utility. In contrast, detection-based approaches are more efficient and practical for deployment in real-world applications. However, the fundamental distinctions between adversarial and benign prompts remain poorly understood. In this work, we introduce CurvaLID, a novel defence framework that efficiently detects adversarial prompts by leveraging their geometric properties. It is agnostic to the type of LLM, offering a unified detection framework across diverse adversarial prompts and LLM architectures. CurvaLID builds on the geometric analysis of text prompts to uncover their underlying differences. We theoretically extend the concept of curvature via the Whewell equation into an $n$-dimensional word embedding space, enabling us to quantify local geometric properties, including semantic shifts and curvature in the underlying manifolds. To further enhance our solution, we leverage Local Intrinsic Dimensionality (LID) to capture complementary geometric features of text prompts within adversarial subspaces. Our findings show that adversarial prompts exhibit distinct geometric signatures from benign prompts, enabling CurvaLID to achieve near-perfect classification and outperform state-of-the-art detectors in adversarial prompt detection. CurvaLID provides a reliable and efficient safeguard against malicious queries as a model-agnostic method that generalises across multiple LLMs and attack families.

对抗攻击提示检测几何分析LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。