arXiv:2603.07886cs.CLcs.AI2026-03被引 1

评测大模型处理复杂指令的能力,真实还原工业场景中的多层约束与流程控制。

CCR-Bench: A Comprehensive Benchmark for Evaluating LLMs on Complex Constraints, Control Flows, and Real-World Cases

  • 任务设计融合内容与格式的深度耦合,模拟真实复杂需求。
  • 模型在复杂分解、条件推理和流程规划上表现明显不足,最高误差超40%。
  • 专为工业级应用设计,适合评估模型在真实场景中的落地能力。

提升大语言模型(LLMs)理解复杂指令的能力,对其实现真实世界应用至关重要。然而,现有评估方法常将指令复杂性简化为原子约束的简单叠加,未能充分捕捉内容与格式交织、逻辑流程控制及现实应用场景中产生的高维复杂性,导致评估实践与实际需求存在显著差距。为此,我们提出CCR-Bench,一个新型基准,用于评估模型在复杂指令下的遵循能力。其特点包括:(1) 任务说明中内容与格式要求深度耦合;(2) 指令涉及复杂的任务分解、条件推理与过程规划;(3) 所有评估样本均源自真实工业场景。在CCR-Bench上的大量实验表明,即使最先进模型也存在显著性能缺陷,明确量化了当前模型能力与真实世界指令理解需求之间的差距。我们认为,CCR-Bench提供了一个更严格、更真实的评估框架,推动大模型向具备工业级复杂任务理解与执行能力的下一代发展。

原文摘要 · Abstract (English)

Enhancing the ability of large language models (LLMs) to follow complex instructions is critical for their deployment in real-world applications. However, existing evaluation methods often oversimplify instruction complexity as a mere additive combination of atomic constraints, failing to adequately capture the high-dimensional complexity arising from the intricate interplay of content and format, logical workflow control, and real-world applications. This leads to a significant gap between current evaluation practices and practical demands. To bridge this gap, we introduce CCR-Bench, a novel benchmark designed to assess LLMs' adherence to complex instructions. CCR-Bench is characterized by: (1) deep entanglement of content and formatting requirements in task specifications; (2) instructions that involve intricate task decomposition, conditional reasoning, and procedural planning; and (3) evaluation samples derived entirely from real-world industrial scenarios. Extensive experiments on CCR-Bench demonstrate that even state-of-the-art models exhibit substantial performance deficiencies, clearly quantifying the gap between current LLM capabilities and the demands of realworld instruction understanding. We believe that CCR-Bench offers a more rigorous and realistic evaluation framework, advancing the development of LLMs toward the next generation of models capable of understanding and executing complex tasks in industrial applications.

大模型评测复杂指令工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。