评测大模型遵循指令优先级的能力,发现多数模型在冲突指令下表现大幅下降。
IHEval: Evaluating Language Models on Following the Instruction Hierarchy
- 构建9类任务共3538例的基准IHEval,覆盖多层级指令冲突场景。
- 主流模型在指令冲突时性能骤降,最强开源模型仅48%准确率。
- 揭示模型对指令优先级理解不足,适合关注安全可控生成的研究者。
指令层级(instruction hierarchy)通过设定系统消息、用户消息、对话历史和工具输出的优先顺序,对保障语言模型行为的一致性与安全性至关重要。然而,该领域关注度有限,缺乏全面的评估基准。为此,本文提出IHEval,一个包含3,538个样本的新型基准,涵盖九类任务,涉及不同优先级指令的协同与冲突情况。对多个主流语言模型的评估显示,它们难以识别指令优先级。所有被测模型在面对冲突指令时,性能均显著下降,相较于原始指令遵循能力出现明显退化。其中,表现最优的开源模型在解决此类冲突时仅达到48%的准确率。结果表明,未来语言模型的发展亟需针对指令层级理解进行专门优化。
原文摘要 · Abstract (English)
The instruction hierarchy, which establishes a priority order from system messages to user messages, conversation history, and tool outputs, is essential for ensuring consistent and safe behavior in language models (LMs). Despite its importance, this topic receives limited attention, and there is a lack of comprehensive benchmarks for evaluating models' ability to follow the instruction hierarchy. We bridge this gap by introducing IHEval, a novel benchmark comprising 3,538 examples across nine tasks, covering cases where instructions in different priorities either align or conflict. Our evaluation of popular LMs highlights their struggle to recognize instruction priorities. All evaluated models experience a sharp performance decline when facing conflicting instructions, compared to their original instruction-following performance. Moreover, the most competitive open-source model only achieves 48% accuracy in resolving such conflicts. Our results underscore the need for targeted optimization in the future development of LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。