arXiv:2603.04064cs.LGcs.CV2026-03

仅微调不到0.2%参数,就能在多编码器扩散模型上发动有效后门攻击。

Tuning Just Enough: Lightweight Backdoor Attacks on Multi-Encoder Diffusion Models

  • 只训练低秩适配器,冻结预训练编码器权重
  • 少于0.2%参数微调即可实现成功攻击
  • 揭示多编码器模型中被忽视的隐蔽漏洞

随着文本生成图像扩散模型在实际应用中的普及,后门攻击问题日益受到关注。以往基于文本的后门攻击研究主要集中在单个轻量级文本编码器的扩散模型,而近期采用多个大规模文本编码器的模型在此背景下仍缺乏系统分析。由于多编码器引入了大量可训练参数,一个关键问题是:后门攻击能否在该类模型中依然高效且有效?本文以使用三个不同文本编码器的Stable Diffusion 3为例,研究其文本编码器在后门攻击中的作用,定义四类攻击目标并识别每类目标所需的最小编码器集合。基于此,提出多编码器轻量级攻击方法(MELT),仅训练低秩适配器,保持预训练编码器权重冻结。实验表明,仅需调整少于0.2%的总编码器参数,即可在Stable Diffusion 3上实现有效的后门攻击,揭示了多编码器设置下此前未被充分探索的实际攻击风险。

原文摘要 · Abstract (English)

As text-to-image diffusion models become increasingly deployed in real-world applications, concerns about backdoor attacks have gained significant attention. Prior work on text-based backdoor attacks has largely focused on diffusion models conditioned on a single lightweight text encoder. However, more recent diffusion models that incorporate multiple large-scale text encoders remain underexplored in this context. Given the substantially increased number of trainable parameters introduced by multiple text encoders, an important question is whether backdoor attacks can remain both efficient and effective in such settings. In this work, we study Stable Diffusion 3, which uses three distinct text encoders and has not yet been systematically analyzed for text-encoder-based backdoor vulnerabilities. To understand the role of text encoders in backdoor attacks, we define four categories of attack targets and identify the minimal sets of encoders required to achieve effective performance for each attack objective. Based on this, we further propose Multi-Encoder Lightweight aTtacks (MELT), which trains only low-rank adapters while keeping the pretrained text encoder weight frozen. We demonstrate that tuning fewer than 0.2% of the total encoder parameters is sufficient for successful backdoor attacks on Stable Diffusion 3, revealing previously underexplored vulnerabilities in practical attack scenarios in multi-encoder settings.

后门攻击扩散模型轻量级攻击多编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。