无需训练即可同时控制大模型的语言、越狱和简洁性,效果稳定可靠。
Compositional Multilingual and Behavioral Attribute Steering

- 通过叠加不同属性的控制向量,实现多属性无训练联合调控。
- 在合适层与强度下,单属性控制效果可靠,三属性组合部分成功。
- 控制向量在残差流中近似正交,解释了其可组合性。
本研究探讨大型语言模型中语言与行为控制的可组合性。聚焦于语言、越狱和简洁性三种属性,我们考察了四种指令微调模型(来自两个模型家族和两种规模)中,无需训练的属性控制向量是否可通过加法组合保持各自预期效果。结果表明,单属性控制在合适干预层与强度下均可靠,抽象行为(越狱、简洁性)更倾向中间层,语言控制则偏好早期层。当每个向量注入其最佳层时,两属性加法组合可同时实现目标;三属性组合亦部分成功,解决了先前无训练组合研究遗留的不一致问题。进一步分析发现,这些控制向量在残差流中近似正交,与其可组合性一致。
原文摘要 · Abstract (English)
This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we investigate whether additive, training-free composition of attribute steering vectors can preserve the intended steering effect of each attribute, across four instruction-tuned models from two model families and two size scales. We find that single-attribute steering is reliable for all three attributes, but only within an appropriate combination of intervention layer and steering strength, with abstract behaviors (jailbreak, conciseness) favoring middle layers and language favoring earlier layers. We show that additive composition of two attribute vectors succeeds in steering both attributes simultaneously when each is injected at its own best-performing layer, and that this partially extends to three simultaneously composed attributes, addressing an inconsistency left open by prior work on training-free composition. We further analyze the geometric properties of these steering vectors, finding that they are approximately orthogonal in the residual stream, consistent with their compositional behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。