无需标注数据,用隐藏特征实时拦截不安全内容。
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
- 利用稀疏自编码器提取的可解释潜在特征监测风险
- 在多个模型和场景下优于有监督训练的防护方法
- 无需训练,适合快速部署于各类大模型流式场景
大型语言模型在流式场景中的应用日益广泛,传统事后防护方法因无法实时拦截不安全内容而失效。基于令牌级监督训练的流式防护虽能解决此问题,但依赖昂贵标注且易过拟合。本文挑战了流式安全必须依赖令牌级监督训练的范式,指出优秀的事后防护模型本身已在其隐藏表示中编码了令牌级风险信号。为此,提出NExT-Guard——一种无需训练的框架,通过监控稀疏自编码器(SAEs)提取的可解释潜在特征实现流式安全防护。该方法使用公开基础大模型预训练的SAEs,实现灵活、低成本部署,无需令牌级标注。实验表明,NExT-Guard在多种模型、SAE变体及风险场景下均优于基于监督训练的事后与流式防护方法,展现出卓越鲁棒性。结果证明其为实时安全防护的通用、可扩展范式,加速了流式防护的实际落地。
原文摘要 · Abstract (English)
Large language models are increasingly deployed in streaming scenarios, rendering conventional post-hoc safeguards ineffective as they fail to interdict unsafe content in real-time. While streaming safeguards based on token-level supervised training could address this, they necessitate expensive annotations and suffer from severe overfitting. In this work, we challenge the paradigm that streaming safety must rely on token-level supervised training. Instead, it is an inherent capability of well-trained post-hoc safeguards, as they already encode token-level risk signals in hidden representations. Hence, we introduce NExT-Guard, a training-free framework that achieves streaming safeguards by monitoring interpretable latent features from Sparse Autoencoders (SAEs). It uses pretrained SAEs from publicly available base LLMs, enabling flexible, low-cost deployment without token-level supervision. Experimental results show that NExT-Guard outperforms both post-hoc and streaming safeguards based on supervised training, with superior robustness across models, SAE variants, and risk scenarios. These results make NExT-Guard a universal and scalable paradigm for real-time safety, accelerating the practical deployment of streaming safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。