arXiv:2502.18528cs.CRcs.AI2025-02被引 8

用大模型自动攻破SSH系统,5步内完成渗透测试

ARACNE: An LLM-Based Autonomous Shell Pentesting Agent

  • 设计多大模型协同的自动化渗透架构,可执行真实命令
  • 对战自适应防御系统成功率60%,对战CTF挑战成功率达57.58%
  • 适合安全研究者与红队实战人员参考,验证大模型在真实环境中的能力

我们提出ARACNE,一个针对SSH服务的全自主大模型渗透测试代理,可在真实Linux Shell系统上执行命令。引入支持多大模型的新型代理架构。实验表明,ARACNE在对抗自主防御系统ShelLM时达到60%的成功率,在过线绕行(Over The Wire Bandit)CTF挑战中成功率达57.58%,优于现有技术。获胜时,平均行动次数低于5次。结果表明,多大模型协同是提升动作准确性的有效途径。

原文摘要 · Abstract (English)

We introduce ARACNE, a fully autonomous LLM-based pentesting agent tailored for SSH services that can execute commands on real Linux shell systems. Introduces a new agent architecture with multi-LLM model support. Experiments show that ARACNE can reach a 60\% success rate against the autonomous defender ShelLM and a 57.58\% success rate against the Over The Wire Bandit CTF challenges, improving over the state-of-the-art. When winning, the average number of actions taken by the agent to accomplish the goals was less than 5. The results show that the use of multi-LLM is a promising approach to increase accuracy in the actions.

渗透测试大模型SSH安全自动化攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。