arXiv:2601.07835cs.CRcs.CV2026-01被引 1

为安全运维设计的抗注入大模型,防住恶意指令干扰

SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations

  • 用安全感知的规则框架和动态演化机制防御恶意指令
  • 攻击成功率降低94.7%,正常任务准确率仍保持95.1%
  • 适合需要高可靠AI辅助的安全团队使用

大型语言模型已成为安全运营中心的变革性工具,支持日志自动分析、钓鱼邮件分类和恶意软件解释;然而,在对抗性网络安全环境中部署时,模型易受提示注入攻击影响,即恶意指令嵌入安全数据中操控模型行为。本文提出SecureCAI,一种基于宪法AI原则的新型防御框架,融合安全感知防护机制、自适应宪法演化和直接偏好优化以消除不安全响应模式,应对高风险安全场景中传统安全机制对复杂对抗操纵无效的问题。实验表明,与基线模型相比,SecureCAI将攻击成功率降低94.7%,同时在良性安全分析任务上保持95.1%的准确率。该框架通过持续红队反馈循环实现对新兴攻击策略的动态适应,在持续对抗压力下宪法遵从得分超过0.92,为语言模型能力在实际网络安全工作流中的可信集成奠定了基础,填补了当前对抗域中AI安全方法的关键空白。

原文摘要 · Abstract (English)

Large Language Models have emerged as transformative tools for Security Operations Centers, enabling automated log analysis, phishing triage, and malware explanation; however, deployment in adversarial cybersecurity environments exposes critical vulnerabilities to prompt injection attacks where malicious instructions embedded in security artifacts manipulate model behavior. This paper introduces SecureCAI, a novel defense framework extending Constitutional AI principles with security-aware guardrails, adaptive constitution evolution, and Direct Preference Optimization for unlearning unsafe response patterns, addressing the unique challenges of high-stakes security contexts where traditional safety mechanisms prove insufficient against sophisticated adversarial manipulation. Experimental evaluation demonstrates that SecureCAI reduces attack success rates by 94.7% compared to baseline models while maintaining 95.1% accuracy on benign security analysis tasks, with the framework incorporating continuous red-teaming feedback loops enabling dynamic adaptation to emerging attack strategies and achieving constitution adherence scores exceeding 0.92 under sustained adversarial pressure, thereby establishing a foundation for trustworthy integration of language model capabilities into operational cybersecurity workflows and addressing a critical gap in current approaches to AI safety within adversarial domains.

大模型安全提示注入网络安全防御框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。