arXiv:2604.27249cs.CLcs.AI2026-04

复杂指令导致大模型在对抗评估中出现位置坍塌,影响判断可靠性。

Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation

论文配图:Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
图 1 · 摘自论文原文
  • 通过六种指令梯度测试模型响应模式变化
  • 最复杂指令使答案集中于单一位置(99.9%与87.4%)
  • 适合关注模型安全与评估设计的研究者

当被要求在多项选择评估中故意表现不佳时,语言模型是关注题目内容,还是依赖位置捷径?我们通过六种对抗性指令特异性梯度,在2,000个MMLU-Pro项目上对两个指令微调的大模型(Llama-3-8B和Llama-3.1-8B)进行测试。采用分布筛选(响应位置熵)与独立内容参与标准(难度-准确率相关性)联合刻画各条件。结果显示存在三个非单调的响应区间:模糊指令导致适度准确率下降但内容参与保留;标准伪装与能力模仿指令引发位置熵坍塌但仍有部分内容参与;两步式答题主动回避指令导致极端位置坍塌,答案几乎全部集中于单一位置(99.9%与87.4%),且无内容敏感性。该效果在双模型及四个学术领域均复现。分布坍塌与内容参与可共存(两种筛选标准一致性达50%),表明熵基筛选与难度基评估捕捉的是响应有效性的部分独立维度。结果表明,指令复杂度可决定小规模指令微调模型在贪婪解码下对抗合规所用机制是内容感知还是内容盲视。

原文摘要 · Abstract (English)

When instructed to underperform on multiple-choice evaluations, do language models engage with question content or fall back on positional shortcuts? We map the boundary between these regimes using a six-condition adversarial instruction-specificity gradient administered to two instruction-tuned LLMs (Llama-3-8B and Llama-3.1-8B) on 2,000 MMLU-Pro items. Distributional screening (response-position entropy) and an independent content-engagement criterion (difficulty-accuracy correlation) jointly characterise each condition. The gradient reveals three regimes rather than a monotonic transition. Vague adversarial instructions produce moderate accuracy reduction with preserved content engagement. Standard sandbagging and capability-imitation instructions produce positional entropy collapse with partial content engagement. A two-step answer-aware avoidance instruction produces extreme positional collapse, with near-total concentration on a single response position (99.9% and 87.4%) and no measurable content sensitivity. This was the only multi-step instruction tested, and it produced the most extreme shortcut. The attractor position matches each model's content-absent null-prompt default. The effect replicates across both models and four academic domains. Distributional collapse and content engagement can co-occur (50% concordance between screening criteria), indicating that entropy-based screening and difficulty-based content assessment capture partially independent dimensions of response validity. Results suggest that instruction complexity can determine whether adversarial compliance uses content-aware or content-blind mechanisms in small instruction-tuned LLMs under greedy decoding.

大模型评估对抗攻击位置偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。