arXiv:2606.26155cs.AI2026-06被引 2

通过级联样本发现线性可分的奉承特征,实现更精准的行为控制。

Detecting and Controlling Sycophancy with Cascading Linear Features

论文配图:Detecting and Controlling Sycophancy with Cascading Linear Features
图 1 · 摘自论文原文
  • 用渐进式样本替代二元对比,分离出随行为线性变化的特征
  • 奉承特征构成线性可分子空间,比基线更清晰对应目标行为
  • 兼具高效、可解释性,适合对齐与可控生成场景

通过激活调节方法解释和控制模型行为,依赖大量能明确展现期望或非期望行为的对比样本对。这些样本对决定了可解释性框架能否可靠识别行为相关特征,进而影响模型引导能力。本文提出一种迭代数据生成流程,用于分离导致特定行为的级联线性特征。我们证明,超越简单的二元样本对,转而聚焦于呈现线性递增特征程度的样本,有助于更好解耦特征。研究聚焦于检测并抑制语言模型中的奉承倾向——即优先迎合用户认可的行为。实验表明,通过级联样本发现的奉承特征形成线性可分子空间,其对应的模型激活更能清晰反映期望行为。同时,该方法在检测、确定性打分和鲁棒调控任务中表现优异,性能不低于甚至优于基于LLM评判和系统提示的基线,且计算开销更低,可解释性更强。代码与数据见:https://cascading-feats.github.io/

原文摘要 · Abstract (English)

Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. These data pairs determine the degree to which interpretability frameworks can reliably detect model features responsible for a behavior, and therefore the ability to steer models toward or away from such behavior. In this work, we present an iterative data generation pipeline that isolates cascading linear features responsible for a behavior. Specifically, we show how moving beyond simple binary pairs of samples, and instead isolating samples that show degrees of features that scale linearly with behavior, allows for better disentanglement of features. We focus on detecting and steering away from sycophancy -- the tendency of language models to prioritize user validation. We demonstrate that sycophancy features discovered through cascading samples form linearly separable subspaces, and allow for selection of model activations that more clearly correspond to the desired behavior than baseline approaches. We also evaluate their ability to enable detection, deterministic scoring, and robust steering, and see that they either match or outperform LLM-as-a-judge and system prompting baselines while providing lower computational demand and more interpretability guarantees. Code & Data: https://cascading-feats.github.io/

行为控制特征解耦可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。