arXiv:2509.18458cs.CLcs.AI2025-09被引 3

构建可调认知负荷的逻辑谜题基准,精准诊断大模型推理瓶颈

CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density

  • 基于认知负荷理论,独立调控任务难度、干扰项密度和长度
  • 22个顶尖模型测试显示,任务长度是主要制约因素,干扰项呈倒U型影响
  • 适合研究模型推理机制、优化训练数据或设计评估实验的团队

当前大语言模型长上下文推理评测常混杂内在任务复杂度、干扰项干扰和任务长度等关键因素。为实现更精准的失效分析,我们提出CogniLoad,一个基于认知负荷理论(CLT)的新型合成基准。该基准生成自然语言逻辑谜题,其核心维度可独立调节:内在难度(d)控制内在负荷;干扰项与信号比(ρ)调节外在负荷;任务长度(N)作为内化负荷的操作代理。对22个SOTA推理模型的评估揭示了不同的性能敏感性,发现任务长度是主要约束,且对内在复杂度的容忍度各异,干扰项比例呈现倒U型响应。通过系统性地、因子化地控制这些认知负荷维度,CogniLoad提供了一个可复现、可扩展且诊断性强的工具,用于剖析大模型推理局限并指导未来模型发展。

原文摘要 · Abstract (English)

Current benchmarks for long-context reasoning in Large Language Models (LLMs) often blur critical factors like intrinsic task complexity, distractor interference, and task length. To enable more precise failure analysis, we introduce CogniLoad, a novel synthetic benchmark grounded in Cognitive Load Theory (CLT). CogniLoad generates natural-language logic puzzles with independently tunable parameters that reflect CLT's core dimensions: intrinsic difficulty ($d$) controls intrinsic load; distractor-to-signal ratio ($ρ$) regulates extraneous load; and task length ($N$) serves as an operational proxy for conditions demanding germane load. Evaluating 22 SotA reasoning LLMs, CogniLoad reveals distinct performance sensitivities, identifying task length as a dominant constraint and uncovering varied tolerances to intrinsic complexity and U-shaped responses to distractor ratios. By offering systematic, factorial control over these cognitive load dimensions, CogniLoad provides a reproducible, scalable, and diagnostically rich tool for dissecting LLM reasoning limitations and guiding future model development.

推理评测认知负荷合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。