arXiv:2605.20382cs.CLcs.AI2026-05中稿 · the Sci-FM Worksho…

大模型在指令与行为模式冲突时,往往更听从习惯而非指令。

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

论文配图:Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
图 1 · 摘自论文原文
  • 设计冲突对话测试模型在指令与固有模式对抗下的表现。
  • 指令遵循率仅1%到99%,与模型能力评估无关。
  • 输出越多样越抗干扰,推理也难完全避免行为偏差。

语言模型既被训练为遵循指令,又具备强大的模式补全能力。当二者冲突时会发生什么?我们构建了用户指令要求模型以特定方式行为(如始终输出某标记、用特定语言回答或扮演角色)但助手前N轮固定输出另一模式的对话。在13个模型和16种指令下,最多测试50轮,平均指令遵循率介于1%至99%之间,与标准能力基准基本无关。从遵循指令转向跟随模式的转变具有普遍性,但高度依赖模型。抵抗诱导的能力受指令内容影响:若指令符合模型训练时的价值偏好,抵抗更强;输出格式也起关键作用:多标记输出比单标记输出显著更抗干扰。链式思维推理虽能提升鲁棒性,但无法消除敏感性,且可能导致正确推理却输出错误。当被要求预测自身行为时,模型平均准确率达83.5%,但系统性低估自身对诱导压力的抵抗能力。结果表明,即使能力强劲的模型,在诱导压力下指令遵循依然脆弱,输出多样性而非语义理解才是预测鲁棒性的核心因素。

原文摘要 · Abstract (English)

Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We construct conversations in which a user instruction to behave in a target way T (e.g., always output a specific token, answer in a particular language, or adopt a persona) is opposed by N hardcoded assistant turns demonstrating a competing pattern P. We then measure instruction-following (IF) rates in this setting, across 13 models and 16 different instructions, for up to 50 turns. Average instruction-following rates range from 1% to 99% across models, largely uncorrelated with standard capability benchmarks. The transition from instruction-following to pattern-following is universal but highly model-dependent. Robustness is modulated both by instruction content, with models resisting induction longer when instructions align with their trained value priors, and by output format, with diverse multi-token responses proving substantially more resistant than single-token outputs. Chain-of-thought reasoning improves robustness but does not eliminate susceptibility, and can produce dissociation between correct deliberation and incorrect output. When asked to predict their behavior in this setting, models achieve 83.5% accuracy on average but systematically underestimate their own resistance to induction pressure. These results suggest that instruction-following remains brittle under induction pressure even for otherwise capable models, and that output diversity, rather than semantic engagement with the input, is the primary factor predicting robustness.

大模型指令遵循行为偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。