arXiv:2509.13450cs.AIcs.CL2025-09中稿 · ICML

评测大模型语义操控在九大安全维度的表现,发现方法效果高度依赖模型与目标视角的匹配。

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

  • 构建九类安全维度的模块化评测框架,整合多种主流语义操控方法。
  • 不同方法在社会行为上最易退化(最高76%),拒绝生成与道德判断常相互冲突。
  • 适合关注模型安全性、可控性及多目标平衡的研究者与开发者。

我们提出SteeringSafety,一个涵盖九个安全视角、18个数据集的基准评测体系,用于评估表示空间操控方法在大模型中的表现。现有研究多关注操控通用能力,而本工作聚焦拒绝、偏见、幻觉、社会行为、推理、认知完整性及规范判断等安全维度。该基准提供模块化组件,支持DIM、ACE、CAA、PCA、LAT等方法及其最新改进(如条件操控)的统一实现。在Gemma-2-2B、Llama-3.1-8B和Qwen-2.5-7B上的实验表明,强操控效果取决于方法、模型与具体视角的匹配。例如,DIM整体表现稳定,但所有方法均存在显著纠缠现象:提升某一维度性能常大幅影响其他维度。社会行为最脆弱(退化高达76%),拒绝操控(越狱)常损害常识道德判断(最高下降26%),幻觉操控则导致政治倾向不可预测地变化——模型间从右移25%到左移28%不等。结果表明,必须从多安全视角综合评估操控方法。

原文摘要 · Abstract (English)

We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabilities of representation steering, we focus on safety perspectives including refusal, bias, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment. SteeringSafety provides modularized building blocks for state-of-the-art steering methods, enabling unified implementation of DIM, ACE, CAA, PCA, and LAT with recent enhancements such as conditional steering. Results on Gemma-2-2B, Llama-3.1-8B, and Qwen-2.5-7B show that strong steering performance depends on the pairing of method, model, and specific perspective. For instance, DIM is consistently effective, yet all methods exhibit substantial entanglement, where improving effectiveness on one safety perspective often significantly changes performance on others. Social behaviors are most vulnerable (degradation up to 76%), refusal steering (jailbreaking) frequently compromises normative judgment such as commonsense morality (up to 26%), and hallucination steering shifts political views unpredictably across models, ranging from a 25% shift to the right to a 28% shift to the left. These findings show the need to understand steering methods through multiple safety angles rather than a single target behavior.

大模型安全语义操控评测基准可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。