arXiv:2510.15430cs.CVcs.AI2025-10

提出新框架LoD,高效检测未知越狱攻击。

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

  • 从攻击特定学习转为任务特定学习,提升泛化能力。
  • 在多种未知攻击上检测AUROC均更高,效率更优。
  • 适合关注大模型安全与鲁棒性的研究者使用。

尽管经过大量对齐努力,大型视觉语言模型(LVLM)仍易受越狱攻击影响,带来严重安全风险。现有检测方法或学习攻击特异性参数,导致难以泛化至未见攻击;或依赖启发式原则,限制了准确性和效率。为此,我们提出学习检测(LoD)框架,通过将重点从攻击特定学习转向任务特定学习,实现对未知越狱攻击的精准检测。该框架包含多模态安全概念激活向量模块,用于安全导向表征学习;以及安全模式自编码器模块,实现无监督攻击分类。大量实验表明,本方法在多样化的未知攻击上均取得更高的检测AUROC,同时提升效率。代码已公开于https://anonymous.4open.science/r/Learning-to-Detect-51CB。

原文摘要 · Abstract (English)

Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks, posing serious safety risks. To address this, existing detection methods either learn attack-specific parameters, which hinders generalization to unseen attacks, or rely on heuristically sound principles, which limit accuracy and efficiency. To overcome these limitations, we propose Learning to Detect (LoD), a general framework that accurately detects unknown jailbreak attacks by shifting the focus from attack-specific learning to task-specific learning. This framework includes a Multi-modal Safety Concept Activation Vector module for safety-oriented representation learning and a Safety Pattern Auto-Encoder module for unsupervised attack classification. Extensive experiments show that our method achieves consistently higher detection AUROC on diverse unknown attacks while improving efficiency. The code is available at https://anonymous.4open.science/r/Learning-to-Detect-51CB.

模型安全越狱攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。