arXiv:2606.22470cs.AI2026-06

测试大模型在冲突指令下的应对能力,发现冲突类型比模型大小影响更大。

PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs

论文配图:PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs
图 1 · 摘自论文原文
  • 设计新框架PRIME,系统生成三类冲突指令并分类评估响应行为
  • 不同冲突类型导致多种失败模式,模型规模影响不如冲突类型显著
  • 适合关注模型鲁棒性与指令理解真实性的研究者阅读

大型语言模型常面临冲突指令,但现有评测基准多孤立评估元指令,难以揭示模型对冲突指令的处理机制。本文提出框架PRIME(Prompt Resolution under Incompatible Meta-Instructions Evaluation),系统构建响应长度、输出格式和推理过程三类可校准的冲突指令,采用确定性行为分类体系分析五种指令微调的开源大模型在平衡与自然分布两种设置下的表现。结果表明,冲突类型对模型行为的影响大于模型规模,且不同冲突类别呈现出多样化失败模式。研究强调需发展冲突感知能力,指出仅通过孤立约束无法全面评估模型的指令遵循能力。

原文摘要 · Abstract (English)

Large language models (LLMs) often encounter conflicting prompts, although current instruction following benchmarks assess those meta-instructions in isolation, limiting the insights about how models process conflicting instructions. We introduce a framework \textit{PRIME}(\textit{Prompt Resolution under Incompatible Meta-Instructions Evaluation}) to analyze behavior of LLMs when provided with conflicting instructions. \textit{PRIME} purposefully produces calibrated conflicts across response length, output format, and reasoning; classifying model responses with a deterministic behavioral taxonomy. We are evaluating five instruction tuned open weight LLMs in two distinct settings, balanced and naturally distributed. The conclusion we reach upon analysis is that conflict type is more significant in affecting behavior than model scale, and various failure modes across different categories of conflict. Our findings emphasize the value of developing conflict awareness and suggest ability of LLM to follow instructions cannot be assessed through isolated constraints alone.

大模型指令遵循冲突检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。