arXiv:2510.18892cs.CLcs.LG2025-10被引 4

256个大模型指令遵循能力实测,发现常见失败模式。

When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs

  • 用20个精准提示测试不同任务类别的指令遵守情况。
  • 覆盖格式、内容、逻辑顺序等,发现多数模型在复杂指令上失效。
  • 适合评估模型真实能力,非记忆表现,实用性强。

尽管大语言模型广泛部署,但系统性评估其指令遵循能力仍具挑战。本文提出一种精简评估框架,使用20个精心设计的提示,测试256个经验证可用的模型(来自331个OpenRouter可获取模型)在多样化任务类别中的指令遵循表现。实验于2025年10月14日进行,所有模型均经过基础功能验证以避免选择偏差。相较于依赖大量算力的大型基准,本方法兼具全面性与高效性。每个提示聚焦指令遵循的特定方面,如格式合规、内容约束、逻辑顺序和多步任务执行。涵盖主流厂商(OpenAI、Anthropic、Google、Meta、Mistral)及新兴模型(Qwen、DeepSeek、社区模型),提供横向对比分析。结果揭示出一致的失败模式,并识别出特别具挑战性的指令类型。该工作不仅提供实用评估工具,也贡献了当前大模型指令遵循能力最全面的实证研究之一。

原文摘要 · Abstract (English)

Despite widespread deployment of Large Language Models, systematic evaluation of instruction-following capabilities remains challenging. While comprehensive benchmarks exist, focused assessments that quickly diagnose specific instruction adherence patterns are valuable. As newer models may be trained on existing benchmarks, novel evaluation approaches are needed to assess genuine capabilities rather than memorized performance. This paper presents a streamlined evaluation framework using twenty carefully designed prompts to assess LLM instruction-following across diverse task categories. We demonstrate this framework through a large-scale empirical study conducted on October 14, 2025, testing 256 verified working models from 331 available via OpenRouter. To ensure methodological rigor and prevent selection bias, we first verified each model's basic functionality before inclusion. Unlike large-scale benchmarks requiring extensive computational resources, our approach offers a practical diagnostic tool researchers and practitioners can readily apply. Our methodology builds upon verifiable instructions while introducing a compact test suite balancing comprehensiveness with efficiency. Each prompt targets distinct aspects of instruction following, including format compliance, content constraints, logical sequencing, and multi-step task execution. We evaluate models from major providers (OpenAI, Anthropic, Google, Meta, Mistral) and emerging implementations (Qwen, DeepSeek, community models), providing comparative performance analysis. Our findings reveal consistent failure modes and identify specific instruction types posing particular challenges. This work contributes both a practical evaluation tool and one of the most comprehensive empirical analyses of instruction-following capabilities across the contemporary LLM landscape.

指令遵循模型评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。