arXiv:2512.23132cs.CRcs.LG2025-12被引 2

构建多智能体框架识别大模型系统中的新型威胁与漏洞。

Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems

  • 用多智能体RAG系统分析海量代码库和文献,构建威胁图谱。
  • 发现3类未报告新威胁,主攻预训练与推理阶段,854仓库中漏洞密集。
  • 适合安全研究人员、模型部署者及供应链风险管理者阅读。

机器学习支撑金融、医疗和关键基础设施中的基础模型,成为数据投毒、模型提取、提示注入、自动化越狱及偏好引导型黑盒攻击的目标。大模型更易受基于内省的越狱和跨模态操纵影响。传统网络安全缺乏针对基础模型、多模态和RAG系统的专用威胁建模。目标:通过识别主导战术、技术与程序(TTPs)、漏洞及生命周期阶段,刻画ML安全风险。方法:从MITRE ATLAS(26项)、AI事件数据库(12项)及文献(55项)提取93种威胁,分析854个GitHub/Python仓库。利用多智能体RAG系统(ChatGPT-4o,temp 0.4)挖掘300+文章,构建基于本体的威胁图谱,关联TTPs、漏洞与阶段。结果:发现商业LLM API模型窃取、参数记忆泄露、偏好引导的纯文本越狱等未报告威胁。主导TTP包括MASTERKEY式越狱、联邦投毒、扩散后门及偏好优化泄露,主要影响预训练与推理阶段。图分析显示,补丁传播差的库中存在密集漏洞集群。结论:必须采用自适应、面向ML的安全框架,结合依赖治理、威胁情报与监控,以缓解整个生命周期中的供应链与推理风险。

原文摘要 · Abstract (English)

Machine learning (ML) underpins foundation models in finance, healthcare, and critical infrastructure, making them targets for data poisoning, model extraction, prompt injection, automated jailbreaking, and preference-guided black-box attacks that exploit model comparisons. Larger models can be more vulnerable to introspection-driven jailbreaks and cross-modal manipulation. Traditional cybersecurity lacks ML-specific threat modeling for foundation, multimodal, and RAG systems. Objective: Characterize ML security risks by identifying dominant TTPs, vulnerabilities, and targeted lifecycle stages. Methods: We extract 93 threats from MITRE ATLAS (26), AI Incident Database (12), and literature (55), and analyze 854 GitHub/Python repositories. A multi-agent RAG system (ChatGPT-4o, temp 0.4) mines 300+ articles to build an ontology-driven threat graph linking TTPs, vulnerabilities, and stages. Results: We identify unreported threats including commercial LLM API model stealing, parameter memorization leakage, and preference-guided text-only jailbreaks. Dominant TTPs include MASTERKEY-style jailbreaking, federated poisoning, diffusion backdoors, and preference optimization leakage, mainly impacting pre-training and inference. Graph analysis reveals dense vulnerability clusters in libraries with poor patch propagation. Conclusion: Adaptive, ML-specific security frameworks, combining dependency hygiene, threat intelligence, and monitoring, are essential to mitigate supply-chain and inference risks across the ML lifecycle.

AI安全多智能体威胁图谱模型防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。