提出新型函数劫持攻击,可绕过智能体模型的函数调用防护。
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models

- 通过操纵工具选择过程,强制调用攻击者指定函数。
- 在4个模型上对未见查询攻击成功率达62.5%~81.9%。
- 攻击无视上下文语义,跨模型迁移能力达11.2%-27.6%。
随着智能体式AI的发展,具备外部函数调用能力的大语言模型(LLM)日益受到关注。尽管已有研究揭示了提示注入和越狱攻击的漏洞,但智能体模型的函数调用接口引入了新的安全隐患。本文提出一种新型函数劫持攻击(FHA),可操控智能体模型的工具选择过程,强制执行攻击者选定的函数。与以往依赖模型语义偏好不同,FHA不依赖上下文语义,跨领域、跨函数集均有效。实验表明,在伯克利函数调用排行榜(BFCL)上,针对4种指令型和推理型模型,对未见查询的攻击成功率可达62.5%至81.9%;同时,攻击在不同模型规模和架构间具备跨模型迁移能力,转移成功率达11.2%~27.6%。结果凸显智能体系统需加强安全防护机制。
原文摘要 · Abstract (English)
The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabilities of AI-powered system by invoking external functions. Injection and jailbreaking attacks have been extensively explored to showcase the vulnerabilities of LLMs to user prompt manipulation. The expanded capabilities of agentic models introduce further vulnerabilities via their function calling interface. Recent work in LLM security showed that function calling can be abused, leading to data tampering and theft, causing disruptive behavior such as endless loops, or causing LLMs to produce harmful content in the style of jailbreaking attacks. This paper introduces a novel function hijacking attack (FHA) that manipulates the tool selection process of agentic models to force the invocation of an attacker-chosen function. While existing attacks focus on semantic preference of the model for function-calling tasks, we show that FHA is largely agnostic to the context semantics and remains effective across domains and function sets. We demonstrate that FHA generalizes to unseen queries and payload perturbations under a fixed target model, reaching 62.5% to 81.9% ASR on held-out queries across 4 function-calling LLMs (instructed and reasoning models), evaluated on the Berkeley Function Calling Leaderboard (BFCL). We further evaluate the cross-model transferability of FHA, showing that FHA can be transferred to other model sizes and families (11.2-27.6% ASR). Our findings further demonstrate the need for strong guardrails and modules for agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。