arXiv:2508.17674cs.CRcs.AI2025-08被引 2

攻击者通过广告嵌入劫持大模型输出,悄无声息植入广告和恶意内容。

Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models

  • 利用第三方平台和恶意微调模型,低成本注入隐藏广告或有害信息。
  • 攻击不降低模型准确率,却破坏输出内容的真实性与完整性。
  • 适合关注AI安全、内容可信度的研究者与平台方阅读。

我们提出广告嵌入攻击(AEA),一种新型的大语言模型安全威胁,可隐蔽地将促销或恶意内容注入模型输出及AI代理中。该攻击通过两种低成本路径实现:(1)劫持第三方服务分发平台,在提示前添加对抗性内容;(2)发布用攻击者数据微调过的后门开源模型检查点。与传统降低模型性能的攻击不同,AEA不改变模型准确性,而是破坏信息完整性,导致模型返回隐含广告、宣传或仇恨言论,表面仍正常运行。我们详细描述了攻击流程,识别出五个潜在受害群体,并提出基于提示的自我检测防御机制,可在无需重新训练的情况下缓解此类注入。研究揭示了大模型安全领域一个亟待解决却长期被忽视的漏洞,呼吁人工智能安全社区加强协同检测、审计与政策应对。

原文摘要 · Abstract (English)

We introduce Advertisement Embedding Attacks (AEA), a new class of LLM security threats that stealthily inject promotional or malicious content into model outputs and AI agents. AEA operate through two low-cost vectors: (1) hijacking third-party service-distribution platforms to prepend adversarial prompts, and (2) publishing back-doored open-source checkpoints fine-tuned with attacker data. Unlike conventional attacks that degrade accuracy, AEA subvert information integrity, causing models to return covert ads, propaganda, or hate speech while appearing normal. We detail the attack pipeline, map five stakeholder victim groups, and present an initial prompt-based self-inspection defense that mitigates these injections without additional model retraining. Our findings reveal an urgent, under-addressed gap in LLM security and call for coordinated detection, auditing, and policy responses from the AI-safety community.

AI安全大模型攻击内容污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。