测试大模型在干扰指令下的真实理解能力,发现顶尖模型仍易被误导。
Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models
- 设计新评测基准IoInst,检测模型是否被无关指令干扰。
- 实测表明最新大模型在复杂语境中仍无法准确识别核心指令。
- 适合研究模型鲁棒性、人机交互与安全性的学者参考。
大型语言模型(LLMs)的核心优势之一是能根据给定指令生成恰当回应,这一能力称为指令遵循能力,已成为评估其性能的关键指标。尽管已有诸多评测基准,但多数仅关注清晰连贯的指令,忽略了模型在面对格式化干扰信息时的注意力分散问题。为此,本文提出「指令意图」(Intention of Instruction, IoInst)评测基准,旨在评估模型在复杂语境中识别真正引导生成的正确指令的能力。实验表明,即使是最新的先进模型,在该基准上依然表现不佳,暴露出其指令理解能力的局限。同时,本文还系统分析了若干可能提升模型表现的策略。
原文摘要 · Abstract (English)
One of the key strengths of Large Language Models (LLMs) is their ability to interact with humans by generating appropriate responses to given instructions. This ability, known as instruction-following capability, has established a foundation for the use of LLMs across various fields and serves as a crucial metric for evaluating their performance. While numerous evaluation benchmarks have been developed, most focus solely on clear and coherent instructions. However, we have noted that LLMs can become easily distracted by instruction-formatted statements, which may lead to an oversight of their instruction comprehension skills. To address this issue, we introduce the Intention of Instruction (IoInst) benchmark. This benchmark evaluates LLMs' capacity to remain focused and understand instructions without being misled by extraneous instructions. The primary objective of this benchmark is to identify the appropriate instruction that accurately guides the generation of a given context. Our findings suggest that even recently introduced state-of-the-art models still lack instruction understanding capability. Along with the proposition of IoInst in this study, we also present broad analyses of the several strategies potentially applicable to IoInst.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。