通过分析模型内部表征,实时检测微调中的危险行为偏移。
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

- 用线性方向追踪激活空间中的七个对齐相关特征
- 检测到的偏移集中在低维轴上,解释65.5%方差
- 监测器误报率仅2.9%,适合快速部署前筛查
当窄范围微调导致模型在目标任务外表现出危险行为时,会引发突现性错位(EM)。标准训练信号可能无法察觉此类变化,依赖行为评估进行检测成本高昂。本文探讨是否可通过微调过程中的内部表征来检测EM。在四个7-9B规模的开源大语言模型中,我们利用七种对齐相关特征作为激活空间中的线性方向,追踪训练检查点间的表征漂移。结果发现,与EM相关的漂移集中于一个低维轴,解释了65.5%的方差,揭示了该情境下的几何特征。基于此漂移模式构建的轻量级监测器,在未见扰动类型上实现2.2%的漏检率、2.9%的误报率和0.990的AUROC,优于无监督PCA与SAE基线。对两个14B模型的应力测试、更长微调周期及非对齐初始状态的验证,识别出关键部署边界。这些结果表明,特质空间监控可作为LoRA微调期间行为评估的实用补充,且跨显著不同场景部署可能需重新校准。
原文摘要 · Abstract (English)
Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals can miss this shift, making reliable detection costly if it depends on repeated behavioral evaluation. We ask whether emergent misalignment can instead be detected from internal representations during finetuning. Using seven alignment-relevant traits encoded as linear directions in activation space, we track representational drift across training checkpoints in four open-source 7-9B LLMs. EM-relevant drift concentrates on a low-dimensional axis that explains 65.5% of the variance, revealing a geometric signature in the studied regime. A low-overhead monitor built on this drift profile detects dangerous checkpoints with 2.2% false negative rate, 2.9% false positive rate, and 0.990 AUROC on held-out perturbation types, outperforming unsupervised PCA and SAE baselines. Stress tests on two 14B models, longer finetuning runs, and misaligned starting points identify key deployment boundaries. These results position trait-space monitoring as a practical complement to behavioral evaluation for EM detection during LoRA-based finetuning, while showing that deployment across substantially different regimes may require recalibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。