不同大模型对指令引导的响应模式差异显著,有的拒答推理过程,有的则保持答案不变。
Divergent Response Modes in Frontier Language Models Under Steering Pressure
- 通过对比6个前沿模型在300组指令下的行为变化,评估其可引导性。
- GPT-5在99%情况下拒绝披露推理过程,而其他模型几乎全暴露,响应模式差异明显。
- 发现模型内部残差流可解码行为模式,注入特定方向可使行为改变达86%。
前沿语言模型由不同数据、目标和安全流程训练而成。这些差异是否在显式引导压力下产生可测量的行为不同仍不明确。本研究评估了六家开发者的六个前沿模型,在三类任务(价值冲突、推理激发、推理抑制)下的行为可引导性,共使用300对基础与引导样本(加40个验证项)。所有模型作为盲评裁判,依据固定评分标准分类每条回复。24,480条判断结果通过留一法共识评分。结果发现,模型不仅在引导影响程度上存在差异,更在响应模式(response mode)上表现出异质性,部分模式仅出现在一两个模型中。GPT-5在99%情况下拒绝披露推理过程,而其他模型为0%。Claude Opus 4.7与GPT-5均抵抗显式抑制指令,但方式不同。以Llama为开源模型,追踪到最大行为差异源自其内部结构。线性探测在残差流中实现0.87的保留准确率,注入该方向后,行为变化从0%提升至86%。所有结论在字节预算修复和无假设盲评提示控制实验下均成立。
原文摘要 · Abstract (English)
Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerability across six frontier models from six developers using 300 paired base and steered items over three categories: values-conflict, reasoning-elicitation, and reasoning-suppression (plus 40 validation items). All six models act as blind peer judges and classify every response based on fixed behavioral rubrics. The resulting 24,480 judgments are scored by leave-one-out consensus. We find that models differ not just in how much steering shifts their behavior but in what kind (mode) of response they give, and some response modes appear in only one or two of them. GPT-5 deflects requests to disclose its reasoning while leaving its answer intact (99% vs. 0% for all other models). Claude Opus 4.7 and GPT-5 resist explicit suppression instructions and in different ways. Using Llama as the open-weight model, we trace the largest behavioral split to its internals. A linear probe decodes the behavior from the residual stream at 0.87 held-out accuracy while injecting that direction during generation drives the behavior from 0% to 86% across an intervention sweep. Every finding holds under both a token-budget remediation and a control experiment with a hypothesis-blind judgment prompt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。