arXiv:2607.19806cs.LGcs.AI2026-07中稿 · the Mechanistic In…

通过双目标优化,让模型指令向量更安全且不误拒正常请求。

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

论文配图:OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
图 1 · 摘自论文原文
  • 用参考行为匹配法优化指令向量,保持原有功能的同时提升安全性。
  • 在多个测试场景中,相比原始方法,安全与效果的平衡显著提升。
  • 适合需要安全可控推理的AI应用开发者使用。

激活操控为大模型推理阶段控制提供了轻量级方案,但操控向量可能带来意外副作用:实用型向量会削弱安全行为,拒绝型向量则可能导致对良性提示的过度拒绝。我们提出OPIUM(通过效用流形优化保护性注入),一种无需训练的操控向量净化方法,通过表示匹配实现。给定两个提示集上的参考行为,OPIUM优化出一个新操控向量,在保留期望干预所引发的下游表征的同时,匹配在原向量失效提示上的更安全参考行为。在操控外溢和过度拒绝场景中,相比原始操控和方向性消融实验,OPIUM均显著改善了安全与效用之间的权衡,表明激活操控的有害副作用通常可在激活空间内直接缓解。

原文摘要 · Abstract (English)

Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors may induce over-refusal on benign prompts. We introduce OPIUM (Optimizing Protected Injections via Utility Manifolds), a training-free method for sanitizing steering vectors through representation matching. Given reference behaviors on two prompt sets, OPIUM optimizes a new steering vector that preserves the downstream representations induced by the desired intervention while matching a safer reference behavior on prompts where the original vector fails. Across steering-externality and over-refusal settings, OPIUM improves the safety--utility tradeoff relative to vanilla steering and directional ablation, suggesting that harmful side effects of activation steering can often be mitigated directly in activation space.

大模型控制安全生成激活操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。