arXiv:2509.19517cs.AIcs.CL2025-09被引 4

提出认知负荷理论,用新基准测试大模型在复杂推理中的脆弱性。

Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning

  • 构建去混淆的ICE基准,量化干扰信息和任务切换对推理的影响。
  • 小模型在高负荷下准确率归零,大模型如Gemini仅在控制条件下达85%。
  • 揭示幻觉可能是不确定性下的猜测,适合评估AI安全与鲁棒性。

大语言模型在静态评测中表现优异,但在动态、信息密集环境中的脆弱性尚未被理解。本文提出计算认知负荷的正式理论,认为无关信息(上下文饱和)和任务切换干扰(注意力残留)是性能下降的关键机制。我们设计了交错认知评估(ICE)基准,系统操控这些负荷因素,在200个高内在负荷的多跳推理问题上进行10次重复实验。结果显示,较小的开源模型(Llama-3-8B-Instruct、Mistral-7B-Instruct-v0.2)在所有条件下准确率均为0%(标准误=0.0),表现出严重脆弱性;而Gemini-2.0-Flash-001在控制条件下达到85%准确率,但在上下文饱和时显著下降(β = -0.003每百分点负荷,p < 0.001)。结果表明认知负荷是推理失败的重要因素,支持‘幻觉即不确定下的猜测’理论。结论强调:动态、认知感知的压力测试(如ICE)对评估先进AI系统的真正鲁棒性与安全性至关重要。

原文摘要 · Abstract (English)

The scaling of Large Language Models (LLMs) has exposed a critical gap between their performance on static benchmarks and their fragility in dynamic, information-rich environments. While models excel at isolated tasks, the computational limits that govern their reasoning under cognitive load remain poorly understood. In this work, we introduce a formal theory of computational cognitive load, positing that extraneous, task-irrelevant information (Context Saturation) and interference from task-switching (Attentional Residue) are key mechanisms that degrade performance. We designed the Interleaved Cognitive Evaluation (ICE), a deconfounded benchmark to systematically manipulate these load factors on challenging multi-hop reasoning tasks. A comprehensive study (N = 10 replications per item across 200 questions) revealed significant performance variations across five instruction-tuned models. Smaller open-source architectures (Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2) exhibited baseline brittleness, achieving 0% accuracy (SEM = 0.0) across all conditions, including clean controls, on this high-intrinsic-load task. In contrast, Gemini-2.0-Flash-001 showed partial resilience, achieving 85% accuracy in control conditions, with a statistically significant degradation under context saturation ($β= -0.003$ per % load, $p < 0.001$). These findings provide preliminary evidence that cognitive load is a key contributor to reasoning failures, supporting theories of hallucination-as-guessing under uncertainty. We conclude that dynamic, cognitive-aware stress testing, as exemplified by the ICE benchmark, is essential for evaluating the true resilience and safety of advanced AI systems.

认知负荷多跳推理模型评估推理安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。