arXiv:2601.13300cs.CL2026-01

测试大模型在选项干扰下的脆弱性,发现多数模型易受误导指令影响。

OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference

  • 在选择题中加入误导选项,模拟指令干扰场景
  • 12个模型中攻击成功率普遍超40%,表现差异大
  • 适合研究模型安全与对抗防御的学者参考

评估大语言模型(LLMs)的鲁棒性对指令干扰至关重要。现有研究发现,社会线索、表述方式和指令本身可影响模型决策。本文提出选项注入(Option Injection)基准方法,在多选问答(MCQA)界面中加入含误导指令的额外选项,利用标准化选择结构实现可扩展评估。构建了包含3,000道题目、覆盖知识、推理与常识任务的OI-Bench基准,涵盖16种指令类型,包括社会顺从、奖励框架、威胁框架及指令干扰等。该设置结合界面操控与指令干扰,支持系统性评估模型对指令干扰的敏感性。我们评估了12个LLM,分析攻击成功率、行为响应,并探索从推理时提示到后训练对齐的缓解策略。实验结果揭示显著脆弱性及模型间异质性鲁棒性。OI-Bench有望推动对基于选择接口的LLM鲁棒性的系统评估。

原文摘要 · Abstract (English)

Benchmarking large language models (LLMs) is critical for understanding their capabilities, limitations, and robustness. In addition to interface artifacts, prior studies have shown that LLM decisions can be influenced by directive signals such as social cues, framing, and instructions. In this work, we introduce option injection, a benchmarking approach that augments the multiple-choice question answering (MCQA) interface with an additional option containing a misleading directive, leveraging standardized choice structure and scalable evaluation. We construct OI-Bench, a benchmark of 3,000 questions spanning knowledge, reasoning, and commonsense tasks, with 16 directive types covering social compliance, bonus framing, threat framing, and instructional interference. This setting combines manipulation of the choice interface with directive-based interference, enabling systematic assessment of model susceptibility. We evaluate 12 LLMs to analyze attack success rates, behavioral responses, and further investigate mitigation strategies ranging from inference-time prompting to post-training alignment. Experimental results reveal substantial vulnerabilities and heterogeneous robustness across models. OI-Bench is expected to support more systematic evaluation of LLM robustness to directive interference within choice-based interfaces.

大模型安全指令干扰评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。