用反向表示法让模型自我调节诚实性,无需人工标注。
AntiPaSTO: Self-Supervised Honesty Steering via Anti-Parallel Representations
- 通过反平行向量分离表示,实现内部自监督控制
- 仅用800组合成词对,在6个价值轴上5项胜过提示基线
- 保持双向可控性,适用于拒绝生成场景
随着模型能力提升,人类难以验证其输出的真实性。可扩展的控制方法需满足内部性、自监督和跨分布迁移三要素,现有方法无法同时满足。我们提出AntiPaSTO,通过在模板句中插入两组对比词(共800对),构建沿反平行轴分离的表示空间(+1/-1产生相反偏移),并施加一致性约束防止坍塌。训练不依赖偏好标签,仅使用合成数据。在Gemma-3-1B上,AntiPaSTO在DailyDilemmas任务中相比提示基线提升6.9倍的控制F1分数,且在6个测试价值轴中有5个取得领先。初步结果表明,该方法在提示引发拒绝时仍能维持双向控制能力。
原文摘要 · Abstract (English)
As models grow more capable, humans cannot reliably verify what they say. Scalable steering requires methods that are internal, self-supervised, and transfer out-of-distribution; existing methods satisfy some but not all three. We introduce AntiPaSTO, which separates representations along an antiparallel axis (+1/-1 produce opposite shifts), with coherence constraints preventing collapse. Training uses only two contrasting words inserted into template sentences, with no preference labels. When we use 800 such synthetic pairs on Gemma-3-1B, AntiPaSTO beats prompting baselines by 6.9x Steering F1 on DailyDilemmas and wins on 5 of 6 tested value axes. We also find preliminary evidence that it maintains bidirectional control where prompting triggers refusal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。