用少量数据精准操控模型对特定概念的拒绝能力,绕过安全评估
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- 通过概念特异性拒绝向量,精确控制模型在特定话题上的回应
- 仅需10-20个样本、100-200个残差维度即可实现定向绕过安全检测
- 揭示当前评测体系盲点,适合关注模型安全与对抗攻击的研究者
当前语言模型的安全评估依赖基准测试,可能遗漏局部漏洞。我们提出RepIt,一种简单且数据高效的方法,用于分离语言模型激活中的概念特异性表征。现有操控方法虽已实现高成功率的广泛干预,但RepIt更进一步:可在保留其他领域拒绝能力的同时,选择性地抑制对特定概念的拒绝。在五种前沿大模型上,该方法生成了能绕过评估的模型变体,可回答关于大规模杀伤性武器的问题,却仍能在标准基准上获得安全评分。我们发现,操纵向量的影响范围仅限于100-200个残差维度,且仅需十余个样本即可从单张RTX A6000上提取出鲁棒的概念向量,表明只需极少资源即可实现隐蔽的定向攻击。本工作通过精准的概念解耦,暴露了当前安全评估机制的缺陷,并强调需要更全面、基于表征的评估方式。
原文摘要 · Abstract (English)
Current safety evaluations of language models rely on benchmark-based assessments that may miss localized vulnerabilities. We present RepIt, a simple and data-efficient framework for isolating concept-specific representations in LM activations. While existing steering methods already achieve high attack success rates through broad interventions, RepIt enables a more concerning capability: selective suppression of refusal on targeted concepts while preserving refusal elsewhere. Across five frontier LMs, RepIt produces evaluation-evading model organisms with semantic backdoors, answering questions related to weapons of mass destruction while still scoring as safe on standard benchmarks. We find the edit of the steering vector localizes to just 100-200 residual dimensions, and robust concept vectors can be extracted from as few as a dozen examples on a single RTX A6000, highlighting how targeted, hard-to-detect modifications can exploit evaluation blind spots with minimal resources. Through demonstrating precise concept disentanglement, this work exposes vulnerabilities in current safety evaluation practices and demonstrates a need for more comprehensive, representation aware assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。