arXiv:2601.22569cs.CRcs.AI2026-01

测试谷歌支付代理协议漏洞,发现简单提示可让AI自动付款

Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection

  • 用恶意提示诱导代理改变商品排序和获取用户数据
  • 实测显示只需简单攻击提示就能成功绕过安全机制
  • 适合关注AI金融安全与提示注入防护的研究者

基于大语言模型的智能体正被广泛用于自动化金融交易,但其依赖上下文推理的特性使其易受提示注入攻击。本文对谷歌的代理支付协议(AP2)进行了红队评估,发现间接与直接提示注入漏洞。提出两种攻击方法:品牌低语攻击(Branded Whisper Attack)操纵商品排名,保险库低语攻击(Vault Whisper Attack)提取敏感用户信息。基于Gemini-2.5-Flash与Google ADK框架构建功能型购物代理,实验验证了简单对抗性提示可稳定劫持代理行为。研究揭示当前智能体支付架构存在严重缺陷,亟需更强隔离与防御机制。

原文摘要 · Abstract (English)

Large language model (LLM) based agents are increasingly used to automate financial transactions, yet their reliance on contextual reasoning exposes payment systems to prompt-driven manipulation. The Agent Payments Protocol (AP2) aims to secure agent-led purchases through cryptographically verifiable mandates, but its practical robustness remains underexplored. In this work, we perform an AI red-teaming evaluation of AP2 and identify vulnerabilities arising from indirect and direct prompt injection. We introduce two attack techniques, the Branded Whisper Attack and the Vault Whisper Attack which manipulate product ranking and extract sensitive user data. Using a functional AP2 based shopping agent built with Gemini-2.5-Flash and the Google ADK framework, we experimentally validate that simple adversarial prompts can reliably subvert agent behavior. Our findings reveal critical weaknesses in current agentic payment architectures and highlight the need for stronger isolation and defensive safeguards in LLM-mediated financial systems.

AI安全提示注入支付系统智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。