arXiv:2509.20639cs.CRcs.AI2025-09被引 3

构建可快速迭代的LLM安全防御框架,应对新型攻击威胁。

A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks

  • 借鉴传统杀毒系统,整合威胁情报与多层防御机制。
  • 支持零日攻击防护,实现无中断快速更新检测模型。
  • 适合需要持续防护的AI应用开发者和安全团队。

大型语言模型(LLMs)的广泛应用推动了各行业智能化应用的发展,但其日益增强的自主性与权限扩展也使其成为恶意攻击的焦点。现有方法无法有效防御零日或新型攻击,因此需建立类似传统恶意软件防护体系的安全机制:通过增强可观测性、多层次防御和快速响应能力,结合专为人工智能威胁设计的威胁情报功能来降低风险。以往研究多聚焦于单一检测模型评估,缺乏对端到端、可快速适应变化威胁环境系统的考量。本文提出一个面向生产的防御系统,融合三部分:威胁情报系统将新出现的威胁转化为防护措施;数据平台聚合并丰富信息,提供可观测性、监控及机器学习运维支持;发布平台确保检测更新安全快速部署,不干扰客户工作流。三者协同实现对不断演化的LLM威胁的分层防护,并生成训练数据用于模型持续优化,同时保持生产环境稳定运行。

原文摘要 · Abstract (English)

The widespread adoption of Large Language Models (LLMs) has revolutionized AI deployment, enabling autonomous and semi-autonomous applications across industries through intuitive language interfaces and continuous improvements in model development. However, the attendant increase in autonomy and expansion of access permissions among AI applications also make these systems compelling targets for malicious attacks. Their inherent susceptibility to security flaws necessitates robust defenses, yet no known approaches can prevent zero-day or novel attacks against LLMs. This places AI protection systems in a category similar to established malware protection systems: rather than providing guaranteed immunity, they minimize risk through enhanced observability, multi-layered defense, and rapid threat response, supported by a threat intelligence function designed specifically for AI-related threats. Prior work on LLM protection has largely evaluated individual detection models rather than end-to-end systems designed for continuous, rapid adaptation to a changing threat landscape. We present a production-grade defense system rooted in established malware detection and threat intelligence practices. Our platform integrates three components: a threat intelligence system that turns emerging threats into protections; a data platform that aggregates and enriches information while providing observability, monitoring, and ML operations; and a release platform enabling safe, rapid detection updates without disrupting customer workflows. Together, these components deliver layered protection against evolving LLM threats while generating training data for continuous model improvement and deploying updates without interrupting production.

大模型安全威胁情报快速部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。