用多智能体系统自动攻防机器人网络,100%通过测试。
Environment-Grounded Multi-Agent Workflow for Autonomous Penetration Testing
- 构建共享图内存动态跟踪系统状态和漏洞
- 在ROS/ROS2环境下100%完成攻防挑战(5次测试)
- 适合需要可追溯性的安全评估场景
数字基础设施的复杂性和互联性日益增加,亟需可扩展且可靠的安全部署方法。机器人系统作为关键运营技术,是高度联网的网络物理系统,广泛应用于工业自动化、物流和自主服务领域。本文探索利用大语言模型实现机器人环境下的自动化渗透测试。提出一种面向机器人系统的环境感知多智能体架构,执行过程中动态构建共享图结构记忆,记录可观测系统状态,包括网络拓扑、通信通道、漏洞及已尝试的攻击。该机制实现结构化自动化,同时保障全程可追溯性和上下文管理。在专用机器人夺旗赛(ROS/ROS2)中多次迭代测试,系统在5次运行中均成功完成任务,达成100%成功率,显著优于现有文献基准,且满足欧盟人工智能法案对可追溯性与人工监督的要求。
原文摘要 · Abstract (English)
The increasing complexity and interconnectivity of digital infrastructures make scalable and reliable security assessment methods essential. Robotic systems represent a particularly important class of operational technology, as modern robots are highly networked cyber-physical systems deployed in domains such as industrial automation, logistics, and autonomous services. This paper explores the use of large language models for automated penetration testing in robotic environments. We propose an environment-grounded multi-agent architecture tailored to Robotics-based systems. The approach dynamically constructs a shared graph-based memory during execution that captures the observable system state, including network topology, communication channels, vulnerabilities, and attempted exploits. This enables structured automation while maintaining traceability and effective context management throughout the testing process. Evaluated across multiple iterations within a specialized robotics Capture-the-Flag scenario (ROS/ROS2), the system demonstrated high reliability, successfully completing the challenge in 100\% of test runs (n=5). This performance significantly exceeds literature benchmarks while maintaining the traceability and human oversight required by frameworks like the EU AI Act.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。