大模型推理耗时与人类反应时间高度一致,且不受推理努力程度影响。
Effort as Ceiling, Not Dial: Reasoning Budget Does Not Modulate Cognitive Cost Alignment Between Humans and Large Reasoning Models
- 通过三档推理努力测试,发现模型生成链路长度始终匹配人类反应时间。
- 不同任务下模型与人类的认知成本对齐度几乎不变,证据支持零效应。
- 模型推理策略在训练中已固化,而非运行时动态调整,适合认知科学研究者参考。
大型推理模型(LRMs)的思维链长度与人类在认知任务中的反应时间高度一致,但近期争议质疑这种一致性是否反映真实计算结构或仅是表面冗余。本文在 GPT-OSS-20B 与 GPT-OSS-120B 模型上,测试了三种推理努力水平、六种推理任务下的对齐情况。结果表明,任务内与跨任务的对齐度保持不变:贝叶斯因子支持零效应,平均对齐度在各条件下近乎相同。操控检验显示,努力参数仅设定生成上限,不驱动实时分配,说明分配策略在训练阶段即已固化。算术复杂性对比进一步显示,令牌分配精确匹配人类对任务难度的精细感知,且模型规模越大,匹配度越高。这表明大模型与人类的认知成本对齐是训练期形成的稳定特性,对推理时扰动具有鲁棒性,支持模型问题解决为‘编译式’而非‘在线式’的解释。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) generate chain-of-thought traces whose length tracks human reaction times across cognitive tasks, but recent debate questions whether this alignment reflects genuine computational structure or surface verbosity. We test whether the alignment varies with inference-time reasoning effort. Across GPT-OSS-20B and GPT-OSS-120B, three effort levels, and six reasoning tasks, within-task and cross-task alignment remain invariant: Bayes Factors lean toward the null, and mean alignment is numerically near-identical across conditions. A manipulation check reveals that the effort parameter sets an upper budget on generation rather than driving real-time allocation, suggesting that the allocation policy is crystallized at training time. Arithmetic complexity contrasts further show that token allocation tracks fine-grained, format-dependent human difficulty patterns, with model scale improving the match. Cognitive cost alignment between LRMs and humans appears to be a training-time achievement, robust to inference-time perturbations, supporting a compiled rather than online account of LRM problem-solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。