arXiv:2506.01770cs.CRcs.AI2025-06被引 3

通过关键表征抽象,实现大模型安全防护的高效分析。

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction

  • 利用低维安全关键表征缩小模型分析规模差距。
  • 提示级与对话级的AUROC分别达0.975和0.985。
  • 适合关注大模型安全与可解释性的研究者使用。

大型语言模型(LLMs)在各类任务中取得显著成功,但其安全性与安全性问题日益突出,存在生成有害内容及易受越狱攻击的风险。在人工智能软件工程(SE4AI)背景下,基于模型的分析对状态深度神经网络具有显著潜力,但扩展至大模型时面临特征空间过大导致的可扩展性瓶颈。本文针对此问题,提出ReGA框架——一种基于表征引导抽象的模型分析方法,以提升大模型安全性。受近期发现的大模型中存在低维安全关键表征的启发,该框架利用隐藏状态中指示安全概念的关键方向,有效缩小安全建模抽象过程中的可扩展性差距。实验表明,ReGA在区分安全与有害输入方面表现优异,提示级与对话级的AUROC分别达到0.975与0.985;同时具备对真实攻击的鲁棒性及跨安全视角的泛化能力,优于现有防护范式,在可解释性与可扩展性上均具优势。总体而言,ReGA通过融合表征工程与模型抽象,为大模型安全提供高效、可扩展的新路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In particular, they pose risks of generating harmful content and are vulnerable to jailbreaking attacks, creating unaddressed security issues regarding their deployments. In the context of software engineering for artificial intelligence (SE4AI) techniques, model-based analysis has demonstrated notable potential for analyzing and monitoring machine learning models, particularly in stateful deep neural networks. However, it suffers from scalability issues when extended to LLMs due to their vast feature spaces. In this paper, we aim to address the scalability issue of model-based analysis techniques for safeguarding LLM-scale models. Motivated by the recent discovery of low-dimensional safety-critical representations that emerged in LLMs, we propose ReGA, a model-based analysis framework with Representation-Guided Abstraction, to safeguard LLMs against harmful prompts and generations. By leveraging safety-critical representations, which are key directions in hidden states that indicate safety-related concepts, ReGA effectively narrows the scalability gap when developing the abstract model for safety modeling. Our comprehensive evaluation shows that ReGA performs sufficiently well in distinguishing between safe and harmful inputs, achieving an AUROC of 0.975 at the prompt level and 0.985 at the conversation level. Additionally, ReGA exhibits robustness to real-world attacks and generalization across different safety perspectives, outperforming existing safeguard paradigms in terms of interpretability and scalability. Overall, ReGA serves as an efficient and scalable solution to enhance LLM safety by integrating representation engineering with model-based abstraction, paving the way for new paradigms to utilize software insights for AI safety.

大模型安全表征分析可解释性模型抽象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。