arXiv:2504.04070cs.MAcs.AI2025-04被引 7

用专职监督员实时纠偏,提升多智能体系统的安全与稳定

Enforcement Agents: Enhancing Accountability and Resilience in Multi-Agent AI Frameworks

  • 引入专职监督智能体实时监控并纠正其他智能体的异常行为
  • 配置两个监督员时系统成功率提升至26.7%,远超无监督员的0%
  • 适合关注多智能体安全、可扩展性设计的研究者和开发者

随着自主智能体日益强大且广泛应用,确保其行为安全并保持与系统目标对齐变得愈发重要,尤其在多智能体环境中。当前系统通常依赖智能体自我监控或事后修正,缺乏实时监管机制。本文提出执行代理(Enforcement Agent, EA)框架,在环境内嵌入专用监督智能体,实现对其他智能体的实时监控、违规检测及干预纠正。我们在自研无人机仿真环境中评估该框架,共进行90个实验周期,对比0、1和2个EA配置。结果表明:加入EA显著提升系统安全性——无EA时成功率为0.0%,配置1个EA时升至7.4%,配置2个EA时达26.7%。系统还表现出更强的持续运行能力及更高的恶意无人机纠正率。这些发现凸显了轻量级实时监督在增强多智能体系统对齐性与鲁棒性方面的潜力。

原文摘要 · Abstract (English)

As autonomous agents become more powerful and widely used, it is becoming increasingly important to ensure they behave safely and stay aligned with system goals, especially in multi-agent settings. Current systems often rely on agents self-monitoring or correcting issues after the fact, but they lack mechanisms for real-time oversight. This paper introduces the Enforcement Agent (EA) Framework, which embeds dedicated supervisory agents into the environment to monitor others, detect misbehavior, and intervene through real-time correction. We implement this framework in a custom drone simulation and evaluate it across 90 episodes using 0, 1, and 2 EA configurations. Results show that adding EAs significantly improves system safety: success rates rise from 0.0% with no EA to 7.4% with one EA and 26.7% with two EAs. The system also demonstrates increased operational longevity and higher rates of malicious drone reformation. These findings highlight the potential of lightweight, real-time supervision for enhancing alignment and resilience in multi-agent systems.

多智能体安全对齐实时监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。