大模型执行简单指令能力弱,格式变化会大幅影响表现。
The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
- 测试20个大模型在不同标签格式下的指令遵循能力,发现格式改变导致性能下降超30%。
- 去掉选项内容后,除数字标签外模型均无法达到随机猜测水平,说明缺乏原子指令理解能力。
- 三样本示例无法提升鲁棒性,且非数字标签错误频发,适合关注模型真实指令理解的研究者阅读。
指令微调的大语言模型(IT-LLMs)具备强大的零样本推理能力,但其执行简单、自包含指令的能力尚未充分探索,而这一能力是复杂指令遵循的基础。我们对20个IT-LLMs在修改后的MMLU和MMLU-Pro基准上进行评估,系统性地改变选项标签格式(字母、数字、罗马数字),同时保持语义一致,共覆盖四种范式:(1) 有明确指令时,标签格式变化导致性能显著波动(如罗马数字相比数字下降30.45%),暴露指令格式偏差;(2) 无指令时,性能进一步下降(最高-10.84%),标签敏感性增强,凸显显式引导的重要性;(3) 当选项内容被移除后,仅数字标签下模型能超过随机基线,表明对原子指令的遵循能力薄弱;(4) 三样本示例未能带来显著鲁棒性或保真度提升,生成分析显示标签错误持续存在,尤其在非数字格式中。跨模型规模分析显示,大模型虽准确率更高,但在指令遵循上仍不一致。结果揭示当前指令微调范式的不足,强调需发展针对原子指令遵循的评估方法与训练策略。
原文摘要 · Abstract (English)
Instruction-tuned large language models (IT-LLMs) exhibit strong zero-shot reasoning, yet their ability to execute simple, self-contained instructions remains underexplored, despite this being foundational to complex instruction-following. We evaluate 20 IT-LLMs on modified MMLU and MMLU-Pro benchmarks, by systematically varying the format of option labels (alphabetic, numeric, Roman) while keeping their meaning identical under four paradigms, namely: (1) With explicit instructions, label changes cause large performance shifts (e.g., -30.45\% for Roman vs. numeric), revealing instruction-format bias. (2) Without instructions, performance drops further (up to -10.84\%) and label sensitivity intensifies, underscoring the role of explicit guidance. (3) When option contents are removed, models fail random-choice baselines except with numeric labels, suggesting weak adherence to atomic directives. (4) Three-shot exemplars yield no significant gains in robustness or fidelity, and generation analyses show persistent label errors, especially for non-numeric formats. Across model sizes, larger LLMs achieve higher accuracy but remain inconsistent in instruction adherence. These results expose the insufficiencies of current instruction-tuning paradigms and highlight the need for evaluation methods and training strategies that explicitly target atomic instruction-following.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。