通过前瞻推理主动限制大模型代理的工具获取,防止其过度扩张权力导致安全风险。
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning

- 基于环境建模进行前瞻推理,提前识别并过滤高风险工具。
- 在三个基准测试中实现安全与任务能力的平衡,风险下降超80%。
- 适合关注大模型代理安全性的研究人员和系统开发者。
随着大语言模型(LLM)代理越来越多地利用模型上下文协议(MCP)在复杂环境中运行,其动作空间的扩大带来了潜在的危险能力,突显了权力追求的风险。虽然更广的动作空间和更强的环境控制力对任务完成至关重要,但也造成了脆弱的风险面,微小错误或幻觉可能被放大为灾难性失败。为此,我们提出SafeMCP,一种{服务端}防御插件,通过预测未来安全风险来约束工具获取。SafeMCP利用内部世界模型进行前瞻推理,实施两层防御:主动工具过滤以抑制危险的权力扩展,以及即时干预作为故障保护。为训练SafeMCP,我们引入三阶段流程:环境动态建模、安全策略初始化,以及带双重可验证奖励的强化学习(RL)。在PowerSeeking Bench、ToolEmu和AgentHarm上的实验表明,SafeMCP实现了安全均衡,在有效缓解风险的同时保持了代理的任务效用。
原文摘要 · Abstract (English)
As Large Language Model (LLM) agents increasingly leverage the Model Context Protocol (MCP) to operate in complex environments, the expansion of their action spaces offers agents unsafe capabilities and underscores the risk of power-seeking. While broad action space and greater environment influence are essential for task fulfillment, they create a fragile risk surface where minor errors or hallucinations are magnified into catastrophic failures. In response, we propose SafeMCP, a {server-side} defense plugin that constrains tool acquisition via predictive reasoning regarding future safety risks. SafeMCP utilizes an internal world model for look-ahead reasoning to implement a two-tier defense: proactive tool filtering to constrain hazardous power expansion and immediate intervention as a fail-safe. To train SafeMCP, we introduce a three-stage pipeline comprising environmental dynamic grounding, safe policy initialization, and reinforcement learning (RL) with dual verifiable rewards. Experiments on PowerSeeking Bench, ToolEmu, and AgentHarm show that SafeMCP achieves a safe equilibrium, effectively mitigating risks while preserving agent utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。