用大模型理解自然语言指令,生成跨模型高效攻击
DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models

- 让大模型根据自然语言指令生成对抗性扰动
- 10亿参数模型在15个模型上实现有效攻击
- 支持多种攻击类型,适合安全研究者参考
尽管视觉和多模态基础模型支撑着从感知到复杂推理的关键任务,但它们仍极易受到对抗攻击。传统攻击通常局限于单一预设目标,且与特定模型或任务紧密耦合,限制了其在真实场景中的可扩展性和灵活性。本文提出DarkLLM,一种新型攻击框架,通过训练大语言模型将自然语言攻击指令转化为潜在攻击向量,并解码为视觉对抗扰动。借助自然语言指令微调,DarkLLM在单一框架内统一了定向、非定向、分割及多模型攻击,实现了灵活可控的对抗生成,使每条指令都能在异构模型上诱导预期行为。在4项任务、13个数据集和15个模型上的广泛实验表明,仅需10亿参数的DarkLLM即可有效攻击CLIP、SAM及前沿大模型,揭示了现代基础模型的系统性漏洞。
原文摘要 · Abstract (English)
While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, traditional adversarial attacks are typically limited to single, predefined objectives, tightly coupling each attack to a specific model or task, which restricts their scalability and flexibility in real-world scenarios. In this work, we present DarkLLM, a novel attack framework that trains an LLM to translate natural-language attack instructions into latent attack vectors, which are then decoded into visual adversarial perturbations. By leveraging natural-language instruction tuning, DarkLLM not only unifies targeted, untargeted, segmentation, and multi-model attacks within a single framework, but also achieves flexible and controllable adversarial generation, enabling each instruction to produce a perturbation that induces desired behaviors across heterogeneous models. Through extensive experiments across 4 tasks, 13 datasets, and 15 models, we demonstrate that DarkLLM with only 1B parameters can follow attacker instructions and generate highly effective attacks against CLIP, SAM, and frontier LLMs, revealing a systemic vulnerability in modern foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。