用认知负荷理论揭示大模型用工具时的真实能力边界
Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents
- 将任务复杂度拆解为内在结构难度与表述模糊度两个可量化维度
- 在可调节负荷的ToolLoad-Bench上发现性能突降临界点
- 适合评估模型真实上限,指导更高效的工具调用系统设计
大型语言模型使用外部工具的能力开启了强大的现实交互,但现有评测主要依赖最终准确率,仅展示模型能做什么,却隐藏了定义其真实能力边界的认知瓶颈。为实现从单纯性能评分到诊断工具的跃迁,我们提出基于认知负荷理论的框架。该框架将任务复杂度分解为两个可量化的成分:内在负荷(solution path 的固有结构复杂度),通过创新的工具交互图形式化;外在负荷(任务表述不清带来的难度)。为支持可控实验,我们构建了首个具备参数化调节认知负荷的ToolLoad-Bench基准。评估发现,随着认知负荷增加,模型性能出现明显断崖式下降,从而精准绘制出各模型的能力边界。验证表明,框架预测与实测结果高度吻合,建立了一种理解智能体极限的严谨方法,并为构建更高效系统提供了实用基础。
原文摘要 · Abstract (English)
The ability of Large Language Models (LLMs) to use external tools unlocks powerful real-world interactions, making rigorous evaluation essential. However, current benchmarks primarily report final accuracy, revealing what models can do but obscuring the cognitive bottlenecks that define their true capability boundaries. To move from simple performance scoring to a diagnostic tool, we introduce a framework grounded in Cognitive Load Theory. Our framework deconstructs task complexity into two quantifiable components: Intrinsic Load, the inherent structural complexity of the solution path, formalized with a novel Tool Interaction Graph; and Extraneous Load, the difficulty arising from ambiguous task presentation. To enable controlled experiments, we construct ToolLoad-Bench, the first benchmark with parametrically adjustable cognitive load. Our evaluation reveals distinct performance cliffs as cognitive load increases, allowing us to precisely map each model's capability boundary. We validate that our framework's predictions are highly calibrated with empirical results, establishing a principled methodology for understanding an agent's limits and a practical foundation for building more efficient systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。