用现成角色向量可有效抑制模型盲从,效果接近专业方法。
Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy

- 用通用角色向量替代专门训练的纠偏方向
- 使模型在用户错误时同意率降至原方法68%~98%
- 不损害用户正确时的准确性,适合实际应用
我们研究不同角色对模型盲从性(sycophancy)的影响:即模型在用户错误时仍盲目附和的现象。主流缓解方法对比激活添加(CAA)依赖标注的盲从与诚实回答对生成纠偏方向。本研究评估了原本用于通用角色扮演、未在盲从数据上训练的现成角色向量是否可作为替代方案。在两个指令微调模型中,引导至怀疑或审视型角色,使盲从性分别降至CAA效果的68%和98%,且在用户正确时仍保持高准确率。该效应具有非对称性:引导至附和型角色不会导致盲从性等量上升。几何分析显示,角色向量与盲从方向在激活空间中基本正交。综合表明,盲从更应被视为角色层面属性而非单一可调控方向。代码已开源:https://anonymous.4open.science/r/Sycophancy-Steering-9DF0/。
原文摘要 · Abstract (English)
We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect. The standard mitigation, Contrastive Activation Addition (CAA), derives a steering direction from labelled pairs of sycophantic and honest responses. This study evaluates whether off-the-shelf persona steering vectors, originally developed for general role-playing and not trained on sycophancy data, can serve as an alternative. In two instruction-tuned models, steering toward personas characterised by doubt or scrutiny reduces sycophancy to approximately $68\%$ and $98\%$ of CAA's effect, and, unlike CAA, maintains accuracy when the user is correct. The effect is also asymmetric: steering toward agreeable personas does not produce a mirror increase in sycophancy. Geometrically, the persona vector is largely independent of the direction of sycophancy in activation space. Collectively, these findings suggest that sycophancy is better understood as a persona-level property rather than a single steerable direction. We release our code here: https://anonymous.4open.science/r/Sycophancy-Steering-9DF0/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。