构建多智能体LLM系统框架,用于安全任务协同求解。
Specification and Evaluation of Multi-Agent LLM Systems -- Prototype and Cybersecurity Applications
- 设计智能体模式语言,规范多LLM协作结构。
- 在网络安全任务中验证系统可正确完成问答与防护。
- 适合关注智能体协同与安全应用的研究者。
大型语言模型的最新进展展现出在推理能力上的潜力,如OpenAI和DeepSeek模型所示。为将这些模型应用于文本生成之外的特定领域任务,基于LLM的多智能体系统可通过结合推理技术、代码生成与软件执行,解决复杂问题,尤其适用于多个可能专精的LLM协同工作。然而,尽管对LLM、推理技术及应用已有大量评估,其联合建模与综合应用仍缺乏清晰规范。本文开展探索性研究,提出(1)多智能体系统规范方法,引入一种智能体模式语言;(2)通过多智能体系统架构与原型实现规范的执行与评估。该规范语言、系统架构与原型首次在本研究中提出,基于先前研究成果构建。测试案例涵盖网络安全任务,验证了架构与评估方法的可行性。结果显示,利用OpenAI与DeepSeek的LLM,智能体能正确完成问答、服务器安全与网络安全部署任务。
原文摘要 · Abstract (English)
Recent advancements in LLMs indicate potential for novel applications, as evidenced by the reasoning capabilities in the latest OpenAI and DeepSeek models. To apply these models to domain-specific applications beyond text generation, LLM-based multi-agent systems can be utilized to solve complex tasks, particularly by combining reasoning techniques, code generation, and software execution across multiple, potentially specialized LLMs. However, while many evaluations are performed on LLMs, reasoning techniques, and applications individually, their joint specification and combined application are not well understood. Defined specifications for multi-agent LLM systems are required to explore their potential and suitability for specific applications, allowing for systematic evaluations of LLMs, reasoning techniques, and related aspects. This paper reports the results of exploratory research on (1.) multi-agent specification by introducing an agent schema language and (2.) the execution and evaluation of the specifications through a multi-agent system architecture and prototype. The specification language, system architecture, and prototype are first presented in this work, building on an LLM system from prior research. Test cases involving cybersecurity tasks indicate the feasibility of the architecture and evaluation approach. As a result, evaluations could be demonstrated for question answering, server security, and network security tasks completed correctly by agents with LLMs from OpenAI and DeepSeek.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。