arXiv:2503.16248cs.CRcs.AI2025-03被引 28

AI代理在区块链中因记忆被篡改,可能遭致命攻击导致资产损失。

Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents

  • 利用上下文操纵漏洞,恶意注入历史记录或提示
  • 内存注入攻击成功率远超传统提示注入
  • 适合关注Web3安全与AI代理风险的研究者

集成Web3的AI代理虽具自治性与开放性,但与金融协议和不可篡改智能合约交互时存在安全风险。本文研究真实场景下区块链金融生态中AI代理的脆弱性,提出上下文操纵这一综合攻击向量,利用未受保护的输入通道、记忆模块和外部数据源进行攻击。该攻击扩展了传统提示注入,揭示更隐蔽持久的内存注入威胁。以代表性的去中心化AI代理框架ElizaOS为例,展示恶意注入提示或历史记录可引发未经授权的资产转移和协议违规,造成重大经济损失。为量化风险,我们构建了面向Web3的基准测试CrAIBench,涵盖150+真实区块链任务(如代币转账、交易、跨链桥接)及500+攻击测试用例。评估表明,模型对内存注入的脆弱性显著高于提示注入。进一步评估防御方案发现,仅依赖提示注入防护和检测无法有效抵御上下文污染,而微调类防御能显著降低攻击成功率且保持单步任务性能。结果凸显在区块链环境中实现安全且可信的AI代理的紧迫性。

原文摘要 · Abstract (English)

AI agents integrated with Web3 offer autonomy and openness but raise security concerns as they interact with financial protocols and immutable smart contracts. This paper investigates the vulnerabilities of AI agents within blockchain-based financial ecosystems when exposed to adversarial threats in real-world scenarios. We introduce the concept of context manipulation -- a comprehensive attack vector that exploits unprotected context surfaces, including input channels, memory modules, and external data feeds. It expands on traditional prompt injection and reveals a more stealthy and persistent threat: memory injection. Using ElizaOS, a representative decentralized AI agent framework for automated Web3 operations, we showcase that malicious injections into prompts or historical records can trigger unauthorized asset transfers and protocol violations which could be financially devastating in reality. To quantify these risks, we introduce CrAIBench, a Web3-focused benchmark covering 150+ realistic blockchain tasks. such as token transfers, trading, bridges, and cross-chain interactions, and 500+ attack test cases using context manipulation. Our evaluation results confirm that AI models are significantly more vulnerable to memory injection compared to prompt injection. Finally, we evaluate a comprehensive defense roadmap, finding that prompt-injection defenses and detectors only provide limited protection when stored context is corrupted, whereas fine-tuning-based defenses substantially reduce attack success rates while preserving performance on single-step tasks. These results underscore the urgent need for AI agents that are both secure and fiduciarily responsible in blockchain environments.

AI安全Web3内存攻击区块链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。