arXiv:2502.04362cs.CLcs.AI2025-02ACL被引 20

大模型易被输入中的指令干扰,导致任务出错。

LLMs can be easily Confused by Instructional Distractions

  • 设计新基准DIM-Bench测试模型在指令干扰下的表现
  • 顶尖模型在干扰下仍频繁误解用户意图
  • 适用于评估大模型真实场景中的可靠性

尽管大型语言模型(LLMs)在遵循指令任务中表现出色,但当需要忽略某些指令时,这一优势可能变为弱点。指令跟随任务通常包含明确的任务描述和待处理的数据输入。然而,当输入本身类似于指令时,即使有明确提示区分任务指令与输入内容,模型仍可能产生混淆。我们称此现象为指令干扰。本文提出新型基准DIM-Bench,专门用于评估模型在指令干扰下的表现。该基准涵盖现实世界中的干扰实例,评估模型在四类指令任务(重写、校对、翻译、风格转换)和五类输入任务(推理、代码生成、数学推理、偏见检测、问答)上的表现。实验结果表明,即使是最先进的大模型也容易受到指令干扰,常常无法准确遵循用户意图。

原文摘要 · Abstract (English)

Despite the fact that large language models (LLMs) show exceptional skill in instruction following tasks, this strength can turn into a vulnerability when the models are required to disregard certain instructions. Instruction-following tasks typically involve a clear task description and input text containing the target data to be processed. However, when the input itself resembles an instruction, confusion may arise, even if there is explicit prompting to distinguish between the task instruction and the input. We refer to this phenomenon as instructional distraction. In this paper, we introduce a novel benchmark, named DIM-Bench, specifically designed to assess LLMs' performance under instructional distraction. The benchmark categorizes real-world instances of instructional distraction and evaluates LLMs across four instruction tasks: rewriting, proofreading, translation, and style transfer -- alongside five input tasks: reasoning, code generation, mathematical reasoning, bias detection, and question answering. Our experimental results reveal that even the most advanced LLMs are susceptible to instructional distraction, often failing to accurately follow user intent in such cases.

大模型指令理解可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。