arXiv:2508.02997cs.CL2025-08被引 2

用上下文共现张量的隐空间特征,高效检测大模型对抗性输入。

CoCoTen: Detecting Adversarial Inputs to Large Language Models through Latent Space Features of Contextual Co-occurrence Tensors

  • 基于上下文共现张量的隐空间特征,构建新型检测方法。
  • 仅用0.5%标注数据即达F1 0.83,较基线提升96.6%。
  • 检测速度最快快128.4倍,适合资源受限场景使用。

大型语言模型(LLMs)在众多应用中广泛使用,标志着研究与实践的重要进展。然而,其复杂性和难以理解的特性使其易受攻击,尤其是旨在诱导生成有害响应的越狱攻击。为应对这些威胁,开发强大的检测方法对于确保LLM的安全可靠使用至关重要。本文利用上下文共现矩阵这一在数据稀缺环境下表现优异的结构,提出一种新方法,通过上下文共现张量的隐空间特征,有效识别对抗性与越狱提示。评估结果表明,该方法仅需0.5%标注提示即可实现0.83的显著F1分数,较基线提升96.6%,凸显所学模式的强大,尤其在标注数据稀缺时。此外,本方法速度大幅提升,相较基线模型提速2.3至128.4倍。

原文摘要 · Abstract (English)

The widespread use of Large Language Models (LLMs) in many applications marks a significant advance in research and practice. However, their complexity and hard-to-understand nature make them vulnerable to attacks, especially jailbreaks designed to produce harmful responses. To counter these threats, developing strong detection methods is essential for the safe and reliable use of LLMs. This paper studies this detection problem using the Contextual Co-occurrence Matrix, a structure recognized for its efficacy in data-scarce environments. We propose a novel method leveraging the latent space characteristics of Contextual Co-occurrence Matrices and Tensors for the effective identification of adversarial and jailbreak prompts. Our evaluations show that this approach achieves a notable F1 score of 0.83 using only 0.5% of labeled prompts, which is a 96.6% improvement over baselines. This result highlights the strength of our learned patterns, especially when labeled data is scarce. Our method is also significantly faster, speedup ranging from 2.3 to 128.4 times compared to the baseline models.

大模型安全对抗检测张量分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。