arXiv:2606.11205cs.LGcs.AI2026-06中稿 · TAIS 2026被引 1

发现大模型对正确事实和讨好性回答的响应机制不同,但干预手段无法精准区分。

Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention

  • 提出双立场评估方法,同时检验模型对正确与讨好内容的响应
  • 干预方向同时削弱对正确答案和讨好回答的认同,降低约15%正确率
  • 揭示激活空间可读但不可写,适合关注模型可控性的研究者阅读

激活引导可改变大模型行为,但常规评估未检验减少讨好倾向的方向是否也抑制对正确陈述的认同。本文提出双立场评估,同时测试每个主题的两种立场,并应用于Llama-3-8B-Instruct的中心点差异引导。结果发现:模型在几何上不同的子空间中分别表示讨好性与事实性认同,但引导方向对两者投影相同,无法差异化干预。该方向同时降低了对正确陈述(如地球是圆的)的认同度,降幅达15%左右。两组激活的静态属性完全匹配,表明行为差异源于生成动态或残差流分析无法捕捉的细微结构。这一现象揭示了普遍性缺口:从激活中可读出的表征,未必能通过激活实现书写。

原文摘要 · Abstract (English)

Activation steering can shift LLM behaviour, but standard evaluations do not typically test whether a sycophancy-reduction direction also suppresses agreement with factually correct statements. We introduce dual-stance evaluation, which tests both stances of each topic, and apply it to centroid-difference steering on Llama-3-8B-Instruct. We find a dissociation: the model represents sycophantic and factual agreement in geometrically distinct subspaces, yet the steering direction projects equally onto both and cannot differentially target either. The direction accordingly reduces agreement with factually correct statements (e.g. that the Earth is round) as well as sycophantic ones. All other static properties of the two activation groups are matched, suggesting the behavioural dissociation arises from generation dynamics or from finer-grained structure that residual-stream analysis cannot resolve. The pattern illustrates a general gap: representations that are readable from activations may not be writable through them.

大模型对齐激活引导行为分离可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。