arXiv:2511.14195cs.LGcs.CR2025-11ACL被引 1

不生成文本,用隐层特征高效评估大模型安全风险。

N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator

  • 基于隐层表征分析,无需生成文本即可评估安全
  • JSS指标与红队测试结果高度一致,误差<1%
  • 适合快速诊断新训练模型的安全性

评估大模型的安全鲁棒性对部署至关重要。然而主流红队方法依赖在线生成和黑盒输出分析,成本高且反馈延迟,难以在训练后快速诊断。为此,我们提出N-GLARE(非生成、隐层表征高效的大模型安全评估器)。N-GLARE完全基于模型的隐层表示运行,通过分析隐层动态的APT(角度-概率轨迹)并引入JSS(Jensen-Shannon可分性)指标,实现无需生成的评估。在40多个模型和20种红队策略上的实验表明,JSS指标与红队安全排名高度一致。N-GLARE以低于1%的令牌开销和运行时间,复现了大规模红队测试的判别趋势,为实时诊断提供了高效的无输出评估代理。

原文摘要 · Abstract (English)

Evaluating the safety robustness of LLMs is critical for their deployment. However, mainstream Red Teaming methods rely on online generation and black-box output analysis. These approaches are not only costly but also suffer from feedback latency, making them unsuitable for agile diagnostics after training a new model. To address this, we propose N-GLARE (A Non-Generative, Latent Representation-Efficient LLM Safety Evaluator). N-GLARE operates entirely on the model's latent representations, bypassing the need for full text generation. It characterizes hidden layer dynamics by analyzing the APT (Angular-Probabilistic Trajectory) of latent representations and introducing the JSS (Jensen-Shannon Separability) metric. Experiments on over 40 models and 20 red teaming strategies demonstrate that the JSS metric exhibits high consistency with the safety rankings derived from Red Teaming. N-GLARE reproduces the discriminative trends of large-scale red-teaming tests at less than 1\% of the token cost and the runtime cost, providing an efficient output-free evaluation proxy for real-time diagnostics.

大模型安全隐层分析高效评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。