arXiv:2410.13897cs.CRcs.LG2024-10中稿 · NeurIPS被引 3

为生成式AI安全风险提供可动态应对的理论框架

A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation

  • 构建分层监测与实时对抗模拟机制,识别新型攻击路径
  • 发现潜在空间劫持、多模态跨攻击等未被重视的风险
  • 适合研究者和开发者用于设计更安全的生成模型系统

随着生成式AI系统(包括大语言模型和扩散模型)快速发展,其广泛应用带来了传统AI风险评估框架未能涵盖的新且复杂的安全威胁。本文提出一种新颖的正式框架,通过集成自适应实时监控与动态风险缓解策略,对这些新兴安全风险进行分类与抑制。框架识别出此前未被充分关注的风险,包括潜在空间劫持、多模态交叉攻击向量以及反馈回路引发的模型退化。采用分层方法,融合异常检测、持续红队测试与实时对抗模拟以减轻风险。强调形式化验证以保障模型在演化威胁下的鲁棒性与可扩展性。尽管为理论性工作,但建立了详细的方法论与评估指标体系,为未来实证验证奠定基础。该框架填补了当前AI安全领域的空白,为后续研究与实践提供全面路线图。

原文摘要 · Abstract (English)

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment frameworks. This paper introduces a novel formal framework for categorizing and mitigating these emergent security risks by integrating adaptive, real-time monitoring, and dynamic risk mitigation strategies tailored to generative models' unique vulnerabilities. We identify previously under-explored risks, including latent space exploitation, multi-modal cross-attack vectors, and feedback-loop-induced model degradation. Our framework employs a layered approach, incorporating anomaly detection, continuous red-teaming, and real-time adversarial simulation to mitigate these risks. We focus on formal verification methods to ensure model robustness and scalability in the face of evolving threats. Though theoretical, this work sets the stage for future empirical validation by establishing a detailed methodology and metrics for evaluating the performance of risk mitigation strategies in generative AI systems. This framework addresses existing gaps in AI safety, offering a comprehensive road map for future research and implementation.

生成式AI安全风险动态防御形式化验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。