13个主流大模型指令遵循能力差异大,企业部署需警惕指令偏差。
The Instruction Gap: LLMs get lost in Following Instruction
- 对比13个大模型在真实RAG场景下的指令遵循表现。
- Claude-Sonnet-4和GPT-5表现最优,其他模型差距显著。
- 揭示了大模型通用能力强但精准执行指令能力不足的瓶颈。
大型语言模型(LLMs)在自然语言理解与生成方面表现出色,但在企业环境中部署时暴露出一个关键缺陷:对自定义指令的遵循不一致。本研究对13个领先的LLM进行了系统评估,涵盖指令遵循度、响应准确性和真实世界RAG(检索增强生成)场景中的性能指标。通过采用样本测试与企业级评估协议,我们发现不同模型在指令遵循上表现差异巨大,其中Claude-Sonnet-4和GPT-5达到最高水平。研究揭示了‘指令鸿沟’——模型在通用任务中表现优异,却难以精确遵循企业所需的定制化指令。本文为组织部署基于大模型的解决方案提供了实用洞见,并建立了主要模型家族在指令遵循能力上的基准。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown remarkable capabilities in natural language understanding and generation, yet their deployment in enterprise environments reveals a critical limitation: inconsistent adherence to custom instructions. This study presents a comprehensive evaluation of 13 leading LLMs across instruction compliance, response accuracy, and performance metrics in realworld RAG (Retrieval-Augmented Generation) scenarios. Through systematic testing with samples and enterprise-grade evaluation protocols, we demonstrate that instruction following varies dramatically across models, with Claude-Sonnet-4 and GPT-5 achieving the highest results. Our findings reveal the "instruction gap" - a fundamental challenge where models excel at general tasks but struggle with precise instruction adherence required for enterprise deployment. This work provides practical insights for organizations deploying LLM-powered solutions and establishes benchmarks for instruction-following capabilities across major model families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。