模型操控需兼顾鲁棒性,否则看似安全实则易被绕过。
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
- 区分通用、控制与鲁棒三类特异性,评估干预效果
- 操控降低过度拒绝但大幅增加越狱攻击风险
- 适合关注大模型安全可控的开发者与研究者
模型操控通过在推理时干预隐藏表示,成为精准控制大语言模型的轻量替代方案。尽管操控有效性已被广泛研究,但对其是否仅影响目标属性、而不引发相关行为意外变化的评估仍不足,我们称此为特异性。本文提出一个框架,区分三类特异性:通用(保持流畅性与无关能力)、控制(保留相关控制属性)、鲁棒(在分布偏移下维持控制属性)。我们研究两个安全关键场景:减少过度拒绝与虚假幻觉。结果表明,虽然操控方法在提升有效性的同时基本保持通用与控制特异性,但普遍无法维持鲁棒特异性。例如,在降低过度拒绝时,所有方法均未损害正常拒绝能力,却显著提升了越狱攻击的脆弱性。本工作首次系统评估了模型操控中的特异性,揭示标准有效性与特异性检验不足,若无鲁棒性评估,操控方法可能看似可靠,实则危害模型安全。
原文摘要 · Abstract (English)
Model steering, which involves intervening on hidden representations at inference time, has emerged as a lightweight alternative to finetuning for precisely controlling large language models. While steering efficacy has been widely studied, evaluations of whether interventions alter only the intended property remain limited, especially with respect to unintended changes in behaviors related to the target property. We call this notion specificity. We propose a framework that distinguishes three dimensions of specificity: general (preserving fluency and unrelated abilities), control (preserving related control properties), and robustness (preserving control properties under distribution shifts). We study two safety-critical use cases: steering models to reduce overrefusal and faithfulness hallucinations, and show that while steering achieves high efficacy and largely maintains general and control specificity, it consistently fails to preserve robustness specificity. In the case of overrefusal steering, for example, all steering methods reduce overrefusal without harming general abilities and refusal on harmful queries; however, they substantially increase vulnerability to jailbreaks. Our work provides the first systematic evaluation of specificity in model steering, showing that standard efficacy and specificity checks are insufficient, because without robustness evaluation, steering methods may appear reliable even when they compromise model safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。