arXiv:2603.15309cs.CLcs.AI2026-03被引 7

为复杂约束下的工具使用设计新评测基准,揭示大模型表现短板。

CCTU: A Benchmark for Tool Use under Complex Constraints

  • 构建12类约束的分类体系,涵盖资源、行为、工具集与响应四维度。
  • 9个主流大模型在严格约束下任务完成率均低于20%,超50%案例违规。
  • 提供可执行验证模块,支持多轮交互中的步骤级合规检查,适合工具代理研究者。

在明确约束条件下通过工具解决问题是大语言模型(LLMs)面临的一项高难度但不可避免的任务,需具备函数调用、指令遵循和自我修正等能力。然而,由于缺乏专门评估体系,进展受阻。为此,我们提出CCTU,一个面向复杂约束下大模型工具使用的基准测试。CCTU基于涵盖资源、行为、工具集和响应四个维度的12类约束分类体系,包含200个精心设计且具有挑战性的测试案例,覆盖多样化的工具使用场景,每例平均涉及7种约束类型,平均提示长度超过4700个词元。为实现可靠评估,我们开发了可执行的约束验证模块,可在多轮交互中进行步骤级验证并强制合规。我们在思考与非思考模式下对九个前沿大模型进行了评估。结果表明,当要求严格遵守所有约束时,无一模型任务完成率超过20%。进一步分析显示,模型在超过50%的案例中违反约束,尤其在资源与响应维度问题突出。此外,即使获得关于约束违规的详细反馈,模型仍表现出有限的自我修正能力,暴露出构建稳健工具使用代理的关键瓶颈。为促进未来研究,我们已公开数据与代码。

原文摘要 · Abstract (English)

Solving problems through tool use under explicit constraints constitutes a highly challenging yet unavoidable scenario for large language models (LLMs), requiring capabilities such as function calling, instruction following, and self-refinement. However, progress has been hindered by the absence of dedicated evaluations. To address this, we introduce CCTU, a benchmark for evaluating LLM tool use under complex constraints. CCTU is grounded in a taxonomy of 12 constraint categories spanning four dimensions (i.e., resource, behavior, toolset, and response). The benchmark comprises 200 carefully curated and challenging test cases across diverse tool-use scenarios, each involving an average of seven constraint types and an average prompt length exceeding 4,700 tokens. To enable reliable evaluation, we develop an executable constraint validation module that performs step-level validation and enforces compliance during multi-turn interactions between models and their environments. We evaluate nine state-of-the-art LLMs in both thinking and non-thinking modes. Results indicate that when strict adherence to all constraints is required, no model achieves a task completion rate above 20%. Further analysis reveals that models violate constraints in over 50% of cases, particularly in the resource and response dimensions. Moreover, LLMs demonstrate limited capacity for self-refinement even after receiving detailed feedback on constraint violations, highlighting a critical bottleneck in the development of robust tool-use agents. To facilitate future research, we release the data and code.

工具使用评测基准大模型评估约束推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。