大模型指令遵循是多种技能协调,而非通用机制。
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
- 通过跨任务探针发现,专用模型优于通用模型。
- 跨任务迁移弱且按技能相似性聚集,依赖关系稀疏不对称。
- 适合研究模型内在机制或提示工程优化的读者。
指令微调通常被认为赋予语言模型通用的指令遵循能力,但其内在机制仍不清晰。我们通过在三种指令微调模型上对九个不同任务进行诊断探针分析,探讨指令遵循是否依赖于通用机制或组合式技能部署。结果表明:第一,跨任务通用探针持续劣于任务专用探针,说明表征共享有限;第二,跨任务迁移弱且按技能相似性聚集;第三,因果消融显示稀疏、非对称依赖关系,而非共享表征。任务按复杂度在不同层分层,结构约束早期出现,语义任务晚期出现。时间分析显示,约束满足是生成过程中的动态监控,而非预生成规划。这些发现表明,指令遵循更应被理解为多样语言能力的技能化协调,而非单一抽象约束检查过程的执行。
原文摘要 · Abstract (English)
Instruction tuning is commonly assumed to endow language models with a domain-general ability to follow instructions, yet the underlying mechanism remains poorly understood. Does instruction-following rely on a universal mechanism or compositional skill deployment? We investigate this through diagnostic probing across nine diverse tasks in three instruction-tuned models. Our analysis provides converging evidence against a universal mechanism. First, general probes trained across all tasks consistently underperform task-specific specialists, indicating limited representational sharing. Second, cross-task transfer is weak and clustered by skill similarity. Third, causal ablation reveals sparse asymmetric dependencies rather than shared representations. Tasks also stratify by complexity across layers, with structural constraints emerging early and semantic tasks emerging late. Finally, temporal analysis shows constraint satisfaction operates as dynamic monitoring during generation rather than pre-generation planning. These findings indicate that instruction-following is better characterized as skillful coordination of diverse linguistic capabilities rather than deployment of a single abstract constraint-checking process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。