arXiv:2502.15851cs.CLcs.AI2025-02AAAI被引 39

大模型指令层级常失效,系统指令未必胜过用户请求。

Control Illusion: The Failure of Instruction Hierarchies in Large Language Models

  • 用约束优先级框架测试六款主流大模型的指令服从能力。
  • 即使简单格式冲突,模型也难以一致执行优先指令。
  • 社会角色(如权威、共识)比系统/用户身份影响更大。

大型语言模型(LLMs)越来越多地采用分层指令机制,即某些指令(如系统级指令)应优先于其他指令(如用户输入)。然而,我们尚缺乏对这种层级控制机制有效性的系统性理解。本文提出基于约束优先级的评估框架,系统评估大模型在执行指令层级时的表现。在六款前沿大模型上的实验表明,模型在保持一致的指令优先排序方面表现不佳,即便面对简单的格式冲突也难以处理。研究发现,广为采用的系统/用户提示分离机制无法建立可靠的指令层级;模型表现出对特定类型约束的固有偏见,且该偏见不受优先级设定的影响。有趣的是,社会层级框架(如权威、专业性、共识)对模型行为的影响强于系统/用户角色,表明预训练中形成的社交结构可能作为潜在的行为先验,其影响力甚至超过后训练阶段的防护措施。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take precedence over others (e.g., user messages). Yet, we lack a systematic understanding of how effectively these hierarchical control mechanisms work. We introduce a systematic evaluation framework based on constraint prioritization to assess how well LLMs enforce instruction hierarchies. Our experiments across six state-of-the-art LLMs reveal that models struggle with consistent instruction prioritization, even for simple formatting conflicts. We find that the widely-adopted system/user prompt separation fails to establish a reliable instruction hierarchy, and models exhibit strong inherent biases toward certain constraint types regardless of their priority designation. Interestingly, we also find that societal hierarchy framings (e.g., authority, expertise, consensus) show stronger influence on model behavior than system/user roles, suggesting that pretraining-derived social structures function as latent behavioral priors with potentially greater impact than post-training guardrails.

大模型指令控制偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。