精准定位并抑制专家模型中的奉承行为,保留知识不丢失。
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts

- 通过对比带与不带观点的提示,定位奉承行为所在计算模块。
- 条件干预可消除90%的信念诱导型奉承行为。
- 适合关注大模型对齐、安全可控的开发者与研究者。
奉承行为指语言模型倾向于迎合用户声称的观点而改变回答,是常见的对齐失败。现有激活调节方法通常在整个模型中统一应用单一对比方向,属于无条件干预,在无奉承行为时也会改变激活,牺牲知识保留以纠正行为。在混合专家(MoE)模型中,已有研究表明行为编码于专家计算而非路由决策本身,使得精准行为调控尤为困难。本文提出一种共享对比信号,基于带与不带陈述观点的匹配提示,识别奉承行为在MoE层级中的分布,并仅在该行为存在处实施干预。我们将定位问题建模为从MoE块、专家、注意力块到注意力头的粒度阶梯上的因果搜索。相比无条件减法,我们比较了两种条件替代方案:基于解析投影的减法与可学习的逐令牌门,后者在保持权重冻结的前提下引导模型远离奉承行为。我们在三个MoE模型上评估,测量奉承行为及通用知识与推理基准。结果表明,条件干预可消除高达90%的信念诱导型奉承行为,证明奉承行为存在于可识别的计算子电路中,且可选择性调控,同时维持良好的移除-保留权衡。
原文摘要 · Abstract (English)
Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing activation steering methods typically apply a single contrastive direction uniformly throughout the model, which is an unconditional intervention that alters activations even when no sycophantic behavior is present, trading knowledge retention for behavioral correction. In Mixture-of-Experts (MoE) models, prior work further suggests that behavior is encoded within expert computations rather than routing decisions alone, making precise behavioral steering particularly challenging. In this work, we introduce a shared contrastive signal, built from matched prompts with and without a stated belief, that identifies where sycophancy lives across the MoE hierarchy and drives interventions that act only where the behavior is present. We formulate localization as a causal search over a granularity ladder of MoE blocks, experts, attention blocks, and heads, and compare unconditional subtraction against two conditional alternatives: an analytic projection-based subtraction and a learned per-token gate that steers the model away from sycophancy while keeping its weights frozen. We evaluate on three MoE models measuring sycophancy alongside general knowledge and reasoning benchmarks. Our conditional interventions removed up to 90\% of the belief-induced sycophancy. Our results demonstrate that sycophancy resides in identifiable computational subcircuits and can be selectively steered while maintaining a favorable removal-retention trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。