检测CLIP模型在不同使用接口下的后门暴露风险,发现文本编码器可成攻击载体。
Beyond Native Success: Auditing Deployment-Interface Exposure of CLIP Backdoors

- 构建DIFE框架,统一评估后门在各类接口中的表现
- 发现文本侧中毒能引发检索/重排序等接口的强暴露
- 提出BadTextTower,实现文本控制型攻击且不影响纯视觉使用
对比语言-图像预训练模型广泛应用于特征提取、检索、重排序和选择等下游接口。现有CLIP后门攻击通常仅在原生任务上验证,未明确同一中毒检查点在其他接口中是否仍暴露、削弱或失效。我们提出DIFE(Deployment-Interface Footprint Evaluation)框架,用于审计后门CLIP检查点在各类部署接口中的表现。DIFE通过规范接口组件读出、触发通道、目标事件、参考条件和评估指标,使不同评估可比。还引入有效足迹诊断,识别携带风险的可复用组件或组合,并解释风险转移路径。对复现的CLIP后门进行审计揭示:原生成功并非检查点级风险保证,暴露遵循组件足迹;文本侧中毒不产生文本编码器控制;部分耦合攻击仍受机制限制。该审计揭示现有后门关键缺口:文本编码器本身可成为对抗行为的可复用载体。因此我们提出BadTextTower,实现强文本条件下的检索、重排序和选择暴露,同时保持纯视觉使用几乎无害。
原文摘要 · Abstract (English)
Contrastive Language-Image Pre-training models are widely reused across downstream interfaces, including feature extraction, retrieval, reranking, and selection. Existing CLIP backdoor, however, usually validate attacks on a small attack-native task, leaving unclear whether the same poisoned checkpoint remains exposed, weakens, or becomes not applicable when reused through other interfaces. We introduce DIFE, a Deployment-Interface Footprint Evaluation framework that audits backdoored CLIP checkpoints across deployment interfaces. DIFE makes various evaluations comparable by specifying each interface's component readout, trigger channel, target event, reference condition, and metric. DIFE also introduces effective-footprint diagnosis to identify the reusable CLIP component or component combination that carries exposure and explains where risk transfers. Auditing reproduced CLIP backdoors with DIFE reveals a structured landscape: native success is not a checkpoint-level risk certificate, exposure follows component footprints, text-side poisoning does not yield textual-encoder control, and some coupled attacks remain mechanism-bound. This audit reveals a import gapin existing CLIP backdoors: a textual encoder that itself becomes a reusable carrier of adversarial behavior. We therefore introduce BadTextTower to fill this gap. BadTextTower produces strong text-conditioned retrieval, reranking, and selection exposure while leaving visual-only reuse nearly clean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。