用探测器指导模型去毒,能有效保持检测能力。
Probe-based Fine-tuning for Reducing Toxicity
- 用探测器作为训练信号,引导模型减少毒性行为。
- 训练后探测器准确率仍可恢复至高位,尤其通过重训探测器。
- 偏好学习比分类器方法更利于保留可解释性特征。
基于模型激活训练的探测器可识别欺骗、偏见等难以从输出中发现的不良行为,不仅可用于检测,还可作为训练信号激励良好内部过程。然而,当探测器成为训练目标时,可能因符合目标而失去可靠性(好哈特定律)。本文提出两种基于监督微调和直接偏好优化的方法,在去毒任务中测试对抗探测器训练的效果。为保持探测器准确性,尝试了三种策略:使用探测器集成、保留未参与训练的探测器、训练后重新训练探测器。结果表明,偏好优化在维持探测器可检测性方面优于分类器方法,暗示偏好学习更倾向于保持相关表示而非掩盖。探测器多样性实际收益有限,仅通过训练后重训即可恢复高检测精度。研究显示,探测器驱动训练在特定对齐方法中可行,但若可重训探测器,则集成并无必要。
原文摘要 · Abstract (English)
Probes trained on model activations can detect undesirable behaviors like deception or biases that are difficult to identify from outputs alone. This makes them useful detectors to identify misbehavior. Furthermore, they are also valuable training signals, since they not only reward outputs, but also good internal processes for arriving at that output. However, training against interpretability tools raises a fundamental concern: when a monitor becomes a training target, it may cease to be reliable (Goodhart's Law). We propose two methods for training against probes based on Supervised Fine-tuning and Direct Preference Optimization. We conduct an initial exploration of these methods in a testbed for reducing toxicity and evaluate the amount by which probe accuracy drops when training against them. To retain the accuracy of probe-detectors after training, we attempt (1) to train against an ensemble of probes, (2) retain held-out probes that aren't used for training, and (3) retrain new probes after training. First, probe-based preference optimization unexpectedly preserves probe detectability better than classifier-based methods, suggesting the preference learning objective incentivizes maintaining rather than obfuscating relevant representations. Second, probe diversity provides minimal practical benefit - simply retraining probes after optimization recovers high detection accuracy. Our findings suggest probe-based training can be viable for certain alignment methods, though probe ensembles are largely unnecessary when retraining is feasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。