跨模型转移拒绝行为,无需目标数据即可保持性能。
Universal Refusal Circuits Across LLMs: Cross-Model Transfer via Trajectory Replay and Concept-Basis Reconstruction
- 通过概念基重建实现拒绝路径迁移,跨架构通用。
- 8组模型验证:转移后拒绝率下降,性能无损。
- 适合安全对齐研究者与大模型可解释性开发者。
对齐大模型中的拒绝行为通常被视为模型特有,但我们假设其源于跨模型共享的低维语义电路。为此,我们提出基于概念基重建的轨迹重放框架,可在不同架构(如密集型到MoE)和训练方式间,将捐赠模型的拒绝干预无监督地迁移到目标模型。通过概念指纹对齐层,并利用共享的“概念原子配方”重构拒绝方向,将捐赠模型的删减轨迹映射至目标模型语义空间。为保护模型能力,引入基于权重SVD的稳定性约束,将干预投影至低方差权重子空间以避免副作用。在8组模型对上的评估表明,所转移的配方能持续减弱拒绝行为,同时保持原有性能,有力支持了安全对齐的语义普遍性。
原文摘要 · Abstract (English)
Refusal behavior in aligned LLMs is often viewed as model-specific, yet we hypothesize it stems from a universal, low-dimensional semantic circuit shared across models. To test this, we introduce Trajectory Replay via Concept-Basis Reconstruction, a framework that transfers refusal interventions from donor to target models, spanning diverse architectures (e.g., Dense to MoE) and training regimes, without using target-side refusal supervision. By aligning layers via concept fingerprints and reconstructing refusal directions using a shared ``recipe'' of concept atoms, we map the donor's ablation trajectory into the target's semantic space. To preserve capabilities, we introduce a weight-SVD stability guard that projects interventions away from high-variance weight subspaces to prevent collateral damage. Our evaluation across 8 model pairs confirms that these transferred recipes consistently attenuate refusal while maintaining performance, providing strong evidence for the semantic universality of safety alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。