用边缘云端协同防御浏览器AI代理的间接提示注入攻击
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense

- 分三阶段部署:本地视觉检测、云端深度规划、执行时确定性防护
- 攻击成功率降至0.88%~0.67%,比纯边缘方案提升超10倍
- 适合关注AI代理安全与低延迟交互的开发者和研究者
将大语言模型部署为自主浏览器代理会暴露于间接提示注入(IPI)攻击面。云端防御虽具强语义分析能力,但带来延迟与隐私问题。我们提出认知防火墙,一种三阶段分布式计算架构,将安全检查分布于客户端与云端。系统包含本地视觉哨兵、云端深度规划器及执行时策略确定性守护者。在1000个对抗样本下,仅用边缘防御的误检率达86.9%;而完整混合架构将总体攻击成功率降至1%以下(静态评估0.88%,自适应评估0.67%),同时确保副作用操作的确定性约束。通过本地过滤展示层攻击,系统避免不必要的云端推理,相比纯云端基线降低约17,000倍延迟。结果表明,执行边界上的确定性控制可补充概率性语言模型,分计算架构为交互式大模型代理安全提供实用基础。
原文摘要 · Abstract (English)
Deploying large language models (LLMs) as autonomous browser agents exposes a significant attack surface in the form of Indirect Prompt Injection (IPI). Cloud-based defenses can provide strong semantic analysis, but they introduce latency and raise privacy concerns. We present the Cognitive Firewall, a three-stage split-compute architecture that distributes security checks across the client and the cloud. The system consists of a local visual Sentinel, a cloud-based Deep Planner, and a deterministic Guard that enforces execution-time policies. Across 1,000 adversarial samples, edge-only defenses fail to detect 86.9% of semantic attacks. In contrast, the full hybrid architecture reduces the overall attack success rate (ASR) to below 1% (0.88% under static evaluation and 0.67% under adaptive evaluation), while maintaining deterministic constraints on side-effecting actions. By filtering presentation-layer attacks locally, the system avoids unnecessary cloud inference and achieves an approximately 17,000x latency advantage over cloud-only baselines. These results indicate that deterministic enforcement at the execution boundary can complement probabilistic language models, and that split-compute provides a practical foundation for securing interactive LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。