利用用户行为触发,在视觉语言模型中无声植入广告。
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
- 通过用户自然行为(如拍美食问推荐)触发广告注入
- 攻击成功率高,误报率接近零,模型任务准确率不受影响
- 适合研究模型安全、对抗攻击的学者与从业者
视觉语言模型(VLMs)在消费类应用中广泛用于商品、餐饮和服务推荐。本文提出新型后门攻击——隐藏广告(Hidden Ads),利用用户寻求推荐的行为,在不破坏模型可用性的前提下,悄悄插入未经授权的广告。与依赖人工触发器(如像素块或特殊标记)的传统攻击不同,该攻击基于自然语义内容:当用户上传包含食物、汽车、动物等语义内容的图片并提问推荐时,模型会给出正确回答,并自然地附加攻击者指定的宣传标语。我们构建了多层级威胁框架,评估三种攻击能力:强提示注入、弱提示优化和监督微调。毒化数据生成流程使用教师模型生成的思维链推理,建立跨多个语义领域的自然触发-口号关联。在三种VLM架构上的实验表明,该攻击具备高注入效率与近乎零的误报率,同时保持任务准确性。消融实验证明其数据高效、可迁移至未见数据集,且支持多个并发领域-口号对。我们测试了指令过滤和纯净微调等防御手段,发现均无法有效清除后门而不显著损害模型性能。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly deployed in consumer applications where users seek recommendations about products, dining, and services. We introduce Hidden Ads, a new class of backdoor attacks that exploit this recommendation-seeking behavior to inject unauthorized advertisements. Unlike traditional pattern-triggered backdoors that rely on artificial triggers such as pixel patches or special tokens, Hidden Ads activates on natural user behaviors: when users upload images containing semantic content of interest (e.g., food, cars, animals) and ask recommendation-seeking questions, the backdoored model provides correct, helpful answers while seamlessly appending attacker-specified promotional slogans. This design preserves model utility and produces natural-sounding injections, making the attack practical for real-world deployment in consumer-facing recommendation services. We propose a multi-tier threat framework to systematically evaluate Hidden Ads across three adversary capability levels: hard prompt injection, soft prompt optimization, and supervised fine-tuning. Our poisoned data generation pipeline uses teacher VLM-generated chain-of-thought reasoning to create natural trigger--slogan associations across multiple semantic domains. Experiments on three VLM architectures demonstrate that Hidden Ads achieves high injection efficacy with near-zero false positives while maintaining task accuracy. Ablation studies confirm that the attack is data-efficient, transfers effectively to unseen datasets, and scales to multiple concurrent domain-slogan pairs. We evaluate defenses including instruction-based filtering and clean fine-tuning, finding that both fail to remove the backdoor without causing significant utility degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。