arXiv:2606.26502cs.AIcs.CL2026-06

发现大模型与人类在难题处理上分歧:知难而行,但继续思考的策略不同。

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

论文配图:Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation
图 1 · 摘自论文原文
  • 区分难题识别与后续思考分配,揭示计算行为差异本质
  • 人类失败尝试更短,模型失败反而用更多推理令牌
  • 不同任务中策略差异明显,适合研究认知与智能决策者对比

大型推理模型(LRMs)在人类耗时更长的问题上投入更多推理令牌,表明其对难度结构具有类似敏感性。这种对齐揭示了哪些问题引发更多深思,但未说明困难状态如何转化为持续计算。本文区分‘难题识别’(问题难度与可观测深思的关系)和‘思考分配’(已识别难度与后续工作量的关系)。在视觉抽象、直觉物理和关系推理三类任务中对比人类与模型数据。在视觉抽象任务中,模型推理轨迹长度与人类难度排序一致;控制问题身份后,人类成功尝试比失败更长;而模型失败尝试使用的推理令牌多于成功尝试。直觉物理任务中呈现相同分配模式。但在关系推理任务中,相反符号模式消失,而人类-模型间差异仍存在,说明任务结构影响分配表现形式。人类动作次数显示更长尝试对应持续参与;模型失败轨迹则在控制轨迹长度后表现出犹豫或回溯迹象。资源理性模型将这些模式解释为从困难状态进一步计算的预期价值。人类与模型可在识别难题上达成一致,但在持续思考策略上存在差异。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) spend more reasoning tokens on problems that take humans longer, suggesting sensitivity to a similar structure of difficulty. That alignment identifies which problems elicit more deliberation and leaves open how a difficult state is converted into continued computation. We distinguish difficulty registration, the relation between problem difficulty and observable deliberation, from deliberation allocation, the relation between registered difficulty and continued work. We test this distinction in matched human and LRM data from visual abstraction, intuitive physics, and relational reasoning. In visual abstraction, LRM trace length tracks the human ordering of problem difficulty. After problem identity is controlled, successful human attempts are longer than failed human attempts. Failed LRM attempts receive more reasoning tokens than successful model attempts. The same allocation difference appears in intuitive physics. In relational reasoning, the opposite-sign pattern disappears, while the item-controlled human-LRM difference remains, showing that task structure changes the observable form of allocation. Human action counts link longer human trials to continued task engagement, while failed LRM traces show task-dependent signs of hesitation or revisiting after trace length is controlled. A resource-rational account relates these patterns to the expected value of further computation from a difficult state. Humans and LRMs can agree about which problems are difficult while applying different policies to continued deliberation.

大模型推理认知机制人类对比任务结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。