首次针对大模型代理设计定向位翻转攻击,揭示其深层安全漏洞。
Targeted Bit-Flip Attacks on LLM-Based Agents
- 构建首个面向大模型代理的定向位翻转攻击框架
- 成功操控代理输出与外部工具调用,攻击成功率显著提升
- 适用于研究大模型系统安全的开发者与安全研究人员
定向位翻转攻击(BFAs)利用硬件故障操纵模型参数,构成重大安全威胁。尽管已有研究聚焦单步推理模型(如图像分类器),但基于大模型的代理系统具有多阶段流水线和外部工具调用能力,带来了新的攻击面,尚未被探索。本文提出Flip-Agent,首个面向大模型代理的定向位翻转攻击框架,能够同时操纵最终输出与工具调用行为。实验表明,Flip-Agent在真实代理任务中显著优于现有靶向位翻转攻击,揭示了大模型代理系统中的关键安全隐患。
原文摘要 · Abstract (English)
Targeted bit-flip attacks (BFAs) exploit hardware faults to manipulate model parameters, posing a significant security threat. While prior work targets single-step inference models (e.g., image classifiers), LLM-based agents with multi-stage pipelines and external tools present new attack surfaces, which remain unexplored. This work introduces Flip-Agent, the first targeted BFA framework for LLM-based agents, manipulating both final outputs and tool invocations. Our experiments show that Flip-Agent significantly outperforms existing targeted BFAs on real-world agent tasks, revealing a critical vulnerability in LLM-based agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。