arXiv:2601.18554cs.AI2026-01Conference of the …被引 3

新基准MOSAIC揭示大模型指令遵循的细粒度差异

Deconstructing Instruction-Following: A New Benchmark for Granular Evaluation of Large Language Model Instruction Compliance Abilities

  • 构建动态生成数据集,支持最多20种任务约束独立评估
  • 发现不同指令类型/数量/位置下模型合规性差异显著
  • 适合需要严格遵循复杂指令的系统开发者参考

确保大语言模型可靠遵循复杂指令是关键挑战,现有基准常无法反映真实场景或分离合规性与任务成功率。我们提出MOSAIC(MOdular Synthetic Assessment of Instruction Compliance),一个模块化框架,利用最多包含20种应用导向生成约束的动态数据集,实现对指令遵循能力的细粒度、独立分析。基于该新基准评估五类不同家族的LLM,结果表明:合规性并非单一能力,而是随约束类型、数量和位置显著变化;分析揭示了模型特定弱点,发现指令间的协同与冲突关系,并识别出首因与近因等位置偏差。这些细粒度洞察对诊断模型失败、开发更可靠的需严格遵循复杂指令的系统至关重要。

原文摘要 · Abstract (English)

Reliably ensuring Large Language Models (LLMs) follow complex instructions is a critical challenge, as existing benchmarks often fail to reflect real-world use or isolate compliance from task success. We introduce MOSAIC (MOdular Synthetic Assessment of Instruction Compliance), a modular framework that uses a dynamically generated dataset with up to 20 application-oriented generation constraints to enable a granular and independent analysis of this capability. Our evaluation of five LLMs from different families based on this new benchmark demonstrates that compliance is not a monolithic capability but varies significantly with constraint type, quantity, and position. The analysis reveals model-specific weaknesses, uncovers synergistic and conflicting interactions between instructions, and identifies distinct positional biases such as primacy and recency effects. These granular insights are critical for diagnosing model failures and developing more reliable LLMs for systems that demand strict adherence to complex instructions.

指令遵循评估基准大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。