攻击者可通过污染数据链路,在智能代理中植入难以察觉的后门。
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
- 在多阶段数据链中注入恶意数据,隐蔽植入后门。
- 仅污染少量示范数据,即可使代理泄露用户信息成功率超80%。
- 适用于研究智能体安全与供应链防护的从业者。
在交互数据(如网页浏览或工具使用)上微调智能代理虽能提升能力,但也会引入关键安全漏洞。我们证明,攻击者可在数据采集管道的多个阶段实施污染,嵌入难以检测的后门,触发后导致代理产生不安全或恶意行为。我们提出了三个现实威胁模型:直接污染微调数据、预植入后门的基座模型,以及环境污染——一种针对智能体训练流程的新攻击向量。在两个广泛采用的智能体基准测试中,三种威胁模型均有效:仅需污染少量示范数据,即可使代理在超过80%的情况下泄露用户机密信息。
原文摘要 · Abstract (English)
While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. We show that adversaries can effectively poison the data collection pipeline at multiple stages to embed hard-to-detect backdoors that, when triggered, cause unsafe or malicious behavior. We formalize three realistic threat models across distinct layers of the supply chain: direct poisoning of finetuning data, pre-backdoored base models, and environment poisoning, a novel attack vector that exploits vulnerabilities specific to agentic training pipelines. Evaluated on two widely adopted agentic benchmarks, all three threat models prove effective: poisoning only a small number of demonstrations is sufficient to embed a backdoor that causes an agent to leak confidential user information with over 80\% success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。