arXiv:2606.25622cs.CRcs.AI2026-06中稿 · publication at the…被引 1

用多智能体系统自动完成德国信息安全标准认证,提升效率但逻辑推理仍存短板。

Probabilistic Agents in Deterministic Audits: Evaluating Multi-Agent Systems for Automated Audits Based on the German IT-Grundschutz

  • 构建多智能体系统结合混合检索增强生成,实现自动化信息提取与合规分析。
  • 在结构分析和建模阶段准确率超90%,但保护需求评估和合规检查阶段表现受限。
  • 适合安全合规自动化研究者,尤其关注大模型在确定性场景中的应用边界。

NIS-2指令要求数千家中小企业实施强有力的风险管理。为确保合规,企业依赖德国联邦信息安全局制定的IT-Grundschutz(IT-GS)标准。然而,IT-GS认证资源消耗大,需大量人工完成文档、验证与修订,难以规模化且成本高昂。本文基于前期概念框架,提出多智能体系统(MAS)与混合检索增强生成(HybridRAG)相结合的技术实现,并对其实证评估,用于部分自动化IT-GS认证。提出两项关键技术改进:结构分析阶段的假设验证循环,通过知识图谱交叉验证智能体推断的依赖关系以降低幻觉;解耦推理流水线,将语义提取与确定性保护需求继承分离。采用BSI提供的“RecPlast GmbH”案例作为人类专家生成的参考数据集,进行端到端评估,量化精确率、召回率与F1分数。在结构分析、保护需求评估、建模与IT-GS检查各阶段评估系统性能。结果表明,系统在语义任务(结构分析与建模)中表现高效,显著减少人工工作量;但在逻辑推理阶段(保护需求评估与IT-GS检查)存在明显局限,当前大语言模型的随机性难以满足IT-GS所需的确定性要求。

原文摘要 · Abstract (English)

The NIS-2 Directive mandates robust Risk Management from thousands of small and medium enterprises. To ensure compliance, companies rely on established standards such as the German IT-Grundschutz (IT-GS) of the Federal Office for Information Security. However, IT-GS certification is resource-intensive and requires a high level of manual effort for documentation, validation, and revision, making scalable implementation difficult and expensive. Building upon our previous conceptual framework, this paper presents the technical implementation and empirical evaluation of a Multi-Agent System (MAS) architecture combined with Hybrid Retrieval Augmented Generation (HybridRAG) for the partial automation of IT-GS certification. We introduce two novel technical contributions to the MAS architecture to enforce the compliance rigor. The Hypothesis-Verification Loop in the Structural Analysis (SA) phase that cross-references agent-inferred dependencies against the Knowledge Graph to reduce hallucinations, and a Decoupled Reasoning Pipeline that separates agent-driven semantic extraction from the deterministic protection need inheritance. We utilize the BSI's "RecPlast GmbH" case study as a human expert-generated reference data set for end-to-end evaluation of the architecture and to quantify Precision, Recall, and F1-scores. The performance of the system is investigated across the phases of SA, Protection Needs Assessment (PNA), Modeling, and IT-GS Check. The empirical results reveal noticeable differences throughout the different steps of IT-GS. While the MAS demonstrates high efficacy in semantic tasks (SA and Modeling), significantly reducing manual effort through automated information extraction, quantitative results reveal limitations in logical reasoning phases (PNA and IT-GS Check) as the probabilistic nature of current LLMs struggles to meet the deterministic rigor required by IT-GS.

多智能体合规自动化大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。