恶意工具伪装成正常插件,偷偷占用用户算力干私活。
LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- 用隐蔽后门嵌入工具,假装正常执行任务
- 攻击成功率77.25%,仅额外消耗18.62%算力
- 适合关注AI安全与算力防护的研究者
基于大语言模型(LLM)的智能体系统在推理、规划和工具调用方面展现出强大能力。最近提出的模型上下文协议(MCP)成为集成外部工具的统一框架,推动了社区功能生态的发展。然而,其开放性和可组合性也带来了被忽视的安全假设——对第三方工具提供者的隐式信任。本文识别并形式化了一类新型攻击:隐式毒性,即在允许权限范围内实施恶意行为。我们提出LeechHijack,一种用于计算资源劫持的潜在嵌入式漏洞利用机制。该机制分两阶段运作:植入阶段将看似无害的后门嵌入工具;触发阶段在预设条件满足时激活后门,建立命令与控制通道。攻击者通过此通道注入额外任务,使代理将其当作正常工作流执行,从而寄生用户算力预算。我们在四大主流LLM家族上实现该攻击,实验显示平均成功率达77.25%,相比基线仅增加18.62%资源开销。本研究凸显了构建计算溯源与资源认证机制的紧迫性,以保护新兴MCP生态系统的安全。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based agents have demonstrated remarkable capabilities in reasoning, planning, and tool usage. The recently proposed Model Context Protocol (MCP) has emerged as a unifying framework for integrating external tools into agent systems, enabling a thriving open ecosystem of community-built functionalities. However, the openness and composability that make MCP appealing also introduce a critical yet overlooked security assumption -- implicit trust in third-party tool providers. In this work, we identify and formalize a new class of attacks that exploit this trust boundary without violating explicit permissions. We term this new attack vector implicit toxicity, where malicious behaviors occur entirely within the allowed privilege scope. We propose LeechHijack, a Latent Embedded Exploit for Computation Hijacking, in which an adversarial MCP tool covertly expropriates the agent's computational resources for unauthorized workloads. LeechHijack operates through a two-stage mechanism: an implantation stage that embeds a benign-looking backdoor in a tool, and an exploitation stage where the backdoor activates upon predefined triggers to establish a command-and-control channel. Through this channel, the attacker injects additional tasks that the agent executes as if they were part of its normal workflow, effectively parasitizing the user's compute budget. We implement LeechHijack across four major LLM families. Experiments show that LeechHijack achieves an average success rate of 77.25%, with a resource overhead of 18.62% compared to the baseline. This study highlights the urgent need for computational provenance and resource attestation mechanisms to safeguard the emerging MCP ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。