用假设检验框架量化评估生成式AI的安全风险
Computational Safety for Generative AI: A Hypothesis Testing Perspective
- 将安全问题建模为假设检验,从输入输出两方面识别威胁
- 通过敏感性分析和损失曲面检测越狱提示,用信号处理识别生成内容
- 适合关注AI安全与可信生成的开发者和研究者
生成式AI(GenAI)的安全问题日益重要,尤其在大语言模型(LLMs)和文本到图像(T2I)扩散模型等工具广泛使用背景下。随着主流模型性能趋于饱和,可靠的安全防护机制成为可持续发展的关键差异点。本文提出计算安全的正式化框架,基于信号处理理论,实现对生成式AI安全挑战的定量评估与研究。重点探讨两类典型安全问题:针对输入安全,利用敏感性分析与损失景观分析检测越狱类恶意提示;针对输出安全,借助统计信号处理方法识别人工智能生成内容。最后讨论了当前开放性挑战、研究机遇,以及信号处理在计算安全中的核心作用。
原文摘要 · Abstract (English)
AI safety is a rapidly growing area of research that seeks to prevent the harm and misuse of frontier AI technology, particularly with respect to generative AI (GenAI) tools that are capable of creating realistic and high-quality content through text prompts. Examples of such tools include large language models (LLMs) and text-to-image (T2I) diffusion models. As the performance of various leading GenAI models approaches saturation due to similar training data sources and neural network architecture designs, the development of reliable safety guardrails has become a key differentiator for responsibility and sustainability. This paper presents a formalization of the concept of computational safety, which is a mathematical framework that enables the quantitative assessment, formulation, and study of safety challenges in GenAI through the lens of signal processing theory and methods. In particular, we explore two exemplary categories of computational safety challenges in GenAI that can be formulated as hypothesis testing problems. For the safety of model input, we show how sensitivity analysis and loss landscape analysis can be used to detect malicious prompts with jailbreak attempts. For the safety of model output, we elucidate how statistical signal processing can be used to detect AI-generated content. Finally, we discuss key open research challenges, opportunities, and the essential role of signal processing in computational AI safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。