揭示大模型在指令服从中的角色冲突机制,解释为何系统指令常被忽视。
Who is In Charge? Dissecting Role Conflicts in Instruction Following
- 通过线性探测发现冲突信号早期编码,系统与社会冲突分属不同子空间。
- 内部冲突检测在系统-用户冲突中更强,但仅社会线索能稳定引导服从。
- 操控实验显示社会线索反而以无角色方式增强指令遵循,暗示对齐方法需更敏感。
大型语言模型应遵循层级指令,即系统提示优先于用户输入,但近期研究发现它们常忽略此规则,却强烈服从权威或共识等社会线索。本文基于大规模数据集,进一步揭示其行为机制:线性探测显示冲突决策信号在早期被编码,系统-用户冲突与社会冲突形成独立子空间;直接对数归因表明系统-用户冲突的内部检测更强,但仅社会线索能实现一致化解;操控实验发现,尽管利用社会线索,其向量仍以无角色方式显著提升指令遵循能力。这些结果解释了系统服从的脆弱性,并强调需要轻量级、层次敏感的对齐方法。
原文摘要 · Abstract (English)
Large language models should follow hierarchical instructions where system prompts override user inputs, yet recent work shows they often ignore this rule while strongly obeying social cues such as authority or consensus. We extend these behavioral findings with mechanistic interpretations on a large-scale dataset. Linear probing shows conflict-decision signals are encoded early, with system-user and social conflicts forming distinct subspaces. Direct Logit Attribution reveals stronger internal conflict detection in system-user cases but consistent resolution only for social cues. Steering experiments show that, despite using social cues, the vectors surprisingly amplify instruction following in a role-agnostic way. Together, these results explain fragile system obedience and underscore the need for lightweight hierarchy-sensitive alignment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。