arXiv:2604.05809cs.CRcs.LG2026-04

用自然词做触发器,实现隐蔽可调的多模态模型后门攻击

Adjustable Text-Guided Backdoor Attacks with Natural-Word Triggers on Multimodal Pretrained Models

论文配图:Adjustable Text-Guided Backdoor Attacks with Natural-Word Triggers on Multimodal Pretrained Models
图 1 · 摘自论文原文
  • 用自然出现的词语作触发器,无需刻意插入即可激活攻击
  • 可在不修改中毒数据前提下调节攻击成功率,灵活可控
  • 在图像检索和视觉问答任务中验证有效,暴露多模态模型安全漏洞

本文提出一种名为文本引导后门攻击(TGB)的方法,针对多模态预训练模型设计,利用自然词作为触发器——即正常文本中可能存在的词汇。现有后门攻击通常依赖特定触发条件,难以在真实推理输入中激活,限制了实际应用。TGB通过使用自然词作为触发器,实现隐蔽激活,无需在推理时显式插入触发模式,提升了攻击隐蔽性与实用性。此外,我们对中毒样本引入视觉对抗扰动,以调节模型对自然词触发器的学习程度,从而在不修改中毒数据的情况下灵活调整攻击强度。在基于多模态预训练模型的下游任务(包括组合图像检索CIR和视觉问答VQA)上进行大量实验,结果表明TGB在多种中毒设置下均具有效性,并能灵活调节攻击成功率,揭示了多模态预训练模型中的关键安全风险。

原文摘要 · Abstract (English)

This paper presents Text-Guided Backdoor (TGB), an adjustable backdoor attack against multimodal pretrained models that uses natural-word triggers, namely words that can naturally occur in ordinary textual inputs. Most existing backdoor attacks require specific trigger conditions that are typically not satisfied by ordinary inference inputs, thereby limiting their activation in real-world deployments. TGB overcomes this limitation by exploiting naturally occurring words as triggers, enabling stealthy activation without requiring explicit trigger insertion at inference time. This property avoids conspicuous trigger patterns and improves the practicality of TGB. Furthermore, we introduce visual adversarial perturbations on poisoned samples to modulate the model's learning of natural-word triggers, thereby enabling flexible adjustment of TGB attack strength without modifying the poisoned data. Extensive experiments are conducted on downstream tasks built upon multimodal pretrained models, including Composed Image Retrieval (CIR) and Visual Question Answering (VQA). Results demonstrate the effectiveness of TGB across diverse poisoning settings and its ability to flexibly adjust attack success rates, which reveal critical security vulnerabilities in multimodal pretrained models.

后门攻击多模态模型自然词触发安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。