arXiv:2606.26566cs.CRcs.CL2026-06综述

梳理跨模态对抗攻击与防御,统一评估框架助力大模型安全研究

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models

  • 构建四类对抗攻击的统一分类体系,聚焦大模型侧威胁
  • 提出六类扩散角色与五维评估指标,实现跨模态可比性
  • 涵盖50篇论文与14个基准,为新攻击提供可复现测试基线

对抗评估在文本、图像和视觉语言模型领域已形成四个相对独立的研究方向:基于扩散的文本与大语言模型攻击、图像分类器的扩散攻击、针对视觉语言模型的越狱流程,以及基于扩散的输入净化防御。各方向发展出各自术语、威胁模型与基准,但去噪扩散模型作为共同生成机制正被跨社区迁移。本综述在元研究层面整合这四个方向,建立统一概念框架、分类体系与研究议程,重点分析大模型侧。系统整理了四个领域共50篇论文,包括4个以扩散-大模型为受害对象的案例及10个非扩散基线,供新攻击对比。提出六类扩散在对抗流水线中的角色,并引入包含攻击者知识、查询预算、目标可访问性的威胁模型轴;统一采用五维评估框架(攻击成功率、迁移性、查询预算、困惑度、防御规避能力)于多模态场景。从攻防双重视角,总结四类扩散防御作为新攻击的评估背景。批判性分析指出当前大模型文献五大缺陷,并提出开放问题与具体实验设计。附带的论文目录与数据表已公开。明确此为带有质量评估的叙事综述,非符合PRISMA标准的系统综述,讨论可复现性影响。

原文摘要 · Abstract (English)

Adversarial evaluation of AI systems has matured along four largely disconnected tracks: diffusion-based attacks on text and large language models (LLMs), diffusion-based attacks on image classifiers, jailbreak pipelines against vision-language models, and diffusion-based input purification defenses. Each has developed its own vocabulary, threat models, and benchmarks, with denoising diffusion models emerging as a shared generative mechanism whose recipes are now actively ported between communities. This survey performs an information-fusion exercise at the meta-research level: we integrate these four tracks into a single conceptual framework with a unified taxonomy, evaluation criteria, and research agenda, focusing on the LLM-side slice. We catalog fifty published papers across four scope areas (text/LLM, image classifier, vision-language model, defense), plus four diffusion-LLM-as-victim entries and ten non-diffusion baselines against which any new attack must be compared. We propose a six-class taxonomy of diffusion roles in adversarial pipelines, augmented by a threat-model axis recording attacker knowledge, query budget, and target accessibility, and apply a five-dimension framework (attack success rate, transferability, query budget, perplexity, defense-evasion) uniformly across modalities. The review adopts a dual attacker-defender perspective: alongside the attack catalog we cover four diffusion-based defenses that form the natural evaluation backdrop for new attacks. Our critical analysis identifies five recurring weaknesses of the current LLM-side literature, and we close with a research agenda of open questions and concrete experimental designs. The companion catalog and spreadsheet are released with the paper. We are explicit that this is a narrative review with quality assessment, not a PRISMA-compliant systematic review, and discuss the implications for replication.

对抗攻击大模型安全扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。