arXiv:2607.18114cs.CLcs.AI2026-07被引 1

揭示大模型对提示诱导偏差的敏感性源于对齐训练,而非预训练。

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

论文配图:How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
图 1 · 摘自论文原文
  • 通过隐藏层方向分析,发现每种偏差有独立线性方向。
  • 对齐模型中偏差方向可解码并干预,恢复正确答案。
  • 该方法可作为去偏工具,保留多数正确回答。

现代大语言模型极易受微小提示变化影响:一句随意提示、错误标注的示例或伪造的助手回复常导致原本正确的答案被改变。本文研究这种易感性(包括奉承行为与相关提示诱导偏差)在模型中的分布。在五种模型家族和七类偏差中,我们从隐藏状态中提取每种偏差的方向,并通过探针检测、留一数据集外迁移和因果干预三种方式验证。结果表明,这种易感性主要由对齐训练塑造,而非预训练:基础预训练模型对偏差反应更弱,其激活值中额外的提示特异性信号也更弱。在对齐模型中,每种偏差具有可解码且可操控的线性方向,沿此方向干预可恢复所有测试家族中的无偏答案。但不同偏差并不共享同一表征:跨偏差重叠程度因模型而异,即使行为相似的偏差也占据不同方向。该干预方法还提供了一个去偏概念验证,能在保留大部分正确答案的同时,修复相当比例的偏差错误。因此,提示诱导偏差应被视为由对齐训练生成的一组因果有效的线性方向,而非单一缺陷。

原文摘要 · Abstract (English)

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leave-one-dataset-out (LODO) transfer, and causal intervention. The susceptibility is largely shaped by alignment tuning rather than pretraining: pretrained base models generally cave much less to these biases, and their activations carry much weaker cue-specific signal beyond question content. Within aligned models, each bias has a coherent linear direction that we can both decode and steer along, recovering the unbiased answer across every family we test. The biases do not collapse into a single shared representation, however: cross-bias overlap is model-specific, and even behaviorally similar biases occupy different directions. The same intervention also provides a proof-of-concept debiasing tool, recovering a meaningful share of bias-induced errors while preserving most correct answers across all instruct families. Cue-induced bias is therefore best understood not as a single flaw in LLMs but as a family of causally effective linear directions that are largely shaped by alignment tuning.

提示偏差对齐训练去偏方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。