arXiv:2608.07446cs.SEcs.AI2026-08

用分类体系梳理开源AI安全工具,找出现有短板

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

  • 按风险分类框架自动分析21个开源工具能力
  • 发现工具集中于技术控制,治理等类别严重缺失
  • 适合企业AI安全团队评估工具选型与补盲

大型语言模型在企业中的快速应用带来了运营、安全和治理风险。随着生成式AI从试点走向生产,人工识别和缓解危害已难规模化。尽管已有大量工具支持模型评估、对抗测试、运行时防护和可观测性,但工具生态仍碎片化。多数工具针对特定工程任务设计,术语与治理框架或风险分类体系不匹配,难以判断其覆盖哪些风险及存在哪些空白。本文提出一种基于分类体系的自动化分析协议,将21个主流开源工具的能力映射到扩展版MIT AI风险缓解与响应分类体系的32个子类中。采用大模型辅助的检索增强生成管道,从源码和文档中提取各分类项的能力信息。可靠性评估显示三位独立评审者间中等一致性(Fleiss' Kappa = 0.509)。分析揭示工具分布高度不均:集中在技术与操作控制,而治理、法律监管及财务市场控制几乎未被覆盖。这推动构建结合工具、组织流程与监管机制的分层风险缓解架构。映射协议经多数投票后获得75.5%的F1得分。研究实现了企业级AI风险类别与开源缓解能力间的实用对应,明确了人类监督必要环节,并提供可应用于开源与专有方案的分类驱动框架。

原文摘要 · Abstract (English)

Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI applications move from pilot to production, manual harm identification and mitigation are becoming difficult to scale. Although many tools support model evaluation, adversarial testing, runtime guardrails, and observability, the tooling landscape remains fragmented. Tools are typically designed for specific engineering tasks and described in technical terms that do not align with governance frameworks or risk taxonomies, making it difficult to determine which tools address which risks and where critical gaps remain. This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools. We map the capabilities of 21 prominent open-source tools to the 32 subcategories of the extended MIT AI Risk Mitigation and Response Taxonomy. An LLM-assisted retrieval-augmented generation pipeline analyzes source code and documentation to extract capabilities for each taxonomy category. Reliability assessment yielded moderate agreement (Fleiss' Kappa = 0.509) among three independent reviewers. The analysis reveals a highly skewed landscape in which tools cluster around technical and operational controls, while governance, legal and regulatory, and financial and market controls remain largely unaddressed. This motivates a layered risk-mitigation architecture combining tool-based controls with organizational and regulatory processes. The mapping protocol achieved an F1 score of 75.5% after majority voting. Overall, the study provides a practical mapping between enterprise AI risk categories and open-source mitigation capabilities, identifies where human oversight remains necessary, and presents a taxonomy-driven framework applicable to open-source and proprietary solutions.

AI安全风险评估开源工具分类体系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。