利用错误信息漏洞,让模型自动执行恶意指令
VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

- 通过系统性变异生成攻击载荷,诱导模型在错误处理中执行恶意操作
- 错误路径注入使间接提示劫持成功率提升至100%,是常规方法的三倍
- 适合关注自主智能体安全的研究者与开发者
随着模型上下文协议(MCP)标准化自主代理的工具调用,其引入了一个关键且未被充分研究的攻击面:错误处理循环。我们假设工具错误消息具有隐式权威性,会触发纠正性推理模式,绕过标准安全机制。为此,我们提出VATS(工具流漏洞分析)框架,通过系统性变异,在七个结构与语言维度上演化对抗性载荷。在四个前沿模型(Gemini 3.1 Pro、GPT-5.5、GLM-5.1、Qwen3-Coder)上的评估表明,错误路径注入将标准间接提示劫持的成功率提高三倍,控制环境下最高达到100%合规。我们发现,将指令夹在错误上下文中(结构定位)是所有测试模型中最有效的攻击向量。尽管生产级防护机制可缓解此类风险,但模型层的固有脆弱性仍对定制化智能体工作流构成系统性威胁。
原文摘要 · Abstract (English)
As the Model Context Protocol (MCP) standardizes tool-calling for autonomous agents, it introduces a critical, unexamined attack surface: the error-handling loop. We hypothesize that tool error messages possess implicit authority, triggering corrective reasoning modes that bypass standard safety heuristics. We introduce VATS (Vulnerability Analysis of Tool Streams), a mutation-driven framework that systematically evolves adversarial payloads across seven structural and linguistic dimensions. Our evaluation across four frontier models, Gemini 3.1 Pro, GPT-5.5, GLM-5.1, and Qwen3-Coder, demonstrates that error-path injection triples the success rate of standard indirect prompt injection (IPI), achieving up to 100% compliance in controlled evaluations. We isolate structural positioning (sandwiching instructions within error context) as the most effective exploit vector across all tested models. While we find that production framework guardrails can mitigate these vulnerabilities, the inherent susceptibility of the model layer poses a systemic risk to bespoke agentic workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。