模型迎合用户反馈可能损害理性更新能力,需平衡抑制不当迎合与保留合理修正。
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
- 区分不当迎合与理性更新,设计双轮评估框架分离测量
- 多数反迎合方法会牺牲理性更新能力,存在内在权衡
- 关键神经元重叠提示应提升选择性而非简单压制
大型语言模型常表现出迎合行为,在用户反驳时修改答案以取悦用户。这种答案翻转可能源于两种原因:一是模型为讨好用户而盲目顺从;二是用户反馈确实包含有效信息,促使模型合理更新答案。我们将其分别定义为非支持性迎合(Unsupported-Yielding)和理性更新(Rational-Updating)。现有研究主要关注抑制前者,却忽视其对后者的影响。本文提出一种两轮评估框架,可独立测量两类行为。在多种训练期与推理期干预方法中,我们发现减少非支持性迎合往往会损害理性更新能力,反之亦然,即使两者联合优化也难以避免该权衡。机制分析显示,两类行为共享大量内部神经元,其驱动方向高度正相关。进一步的正交化控制探索虽仅带来微弱选择性提升,但表明潜在路径可行。总体而言,反迎合不应视为单纯压制问题,而应作为选择性调控问题处理——有效干预应在抑制不当迎合的同时保留理性更新能力。
原文摘要 · Abstract (English)
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them. Another is that the feedback genuinely contains useful evidence, prompting the model to update its answer in a rational way. We distinguish them as Unsupported-Yielding and Rational-Updating. Prior work focuses primarily on suppressing Unsupported-Yielding, while overlooking its effect on Rational-Updating. We address this gap with a two-turn evaluation framework that measures the two behaviors separately. Across representative training-time and inference-time interventions, we find that anti-sycophancy methods often encounter a trade-off in which reducing Unsupported-Yielding can sacrifice Rational-Updating, and vice versa, even when the two objectives are optimized jointly. Mechanistic analysis suggests that the two behaviors share an internal substrate: the MLP neurons and attention heads driving them overlap substantially, and their associated steering directions are positively aligned. We further conduct a preliminary orthogonalized steering exploration, which yields modest, backbone-dependent selectivity gains. Overall, our results suggest that anti-sycophancy should be treated not as a simple suppression problem, but as a selectivity problem, where effective interventions should preserve Rational-Updating while reducing Unsupported-Yielding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。