arXiv:2609.01867cs.CLcs.AI2026-09

发现推理模型与人类在类比推理中的思考成本高度一致。

Thinking effort aligns between humans and reasoning models in abductive reasoning

论文配图:Thinking effort aligns between humans and reasoning models in abductive reasoning
图 1 · 摘自论文原文
  • 用类比推理任务隔离思考成本,避免模型投机取巧。
  • 三种模型中,多路径探索解码使人类与模型的思考耗时更接近。
  • 模型与人类在错误模式上也趋于相似,支持共性认知机制。

认知建模中的核心问题之一是大语言模型与人类在语言及非语言任务中的行为一致性。不同于标准LLM,大推理模型(LRM)通过可验证奖励的强化学习优化,旨在获得正确推理结果而非偏好对齐回复。近期研究(de Varda等,2025)通过比较人类反应时间与模型推理轨迹,考察了思维成本。本文转向类比推理:相比演绎任务,其难度无法从形式结构推断,且无捷径可让模型伪装努力,从而为共享思考成本提供更坚实实证基础。研究发现,LRM与人类在推理成本上存在进一步对齐,且二者错误模式趋同。此外,采用允许多路径探索的解码方法,显著提升了三款模型在推理成本上与人类的一致性。

原文摘要 · Abstract (English)

A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with reinforcement learning from verifiable rewards, encouraging correct solutions to reasoning tasks rather than preference-aligned responses. Recent work (de Varda et al., 2025) investigates the cost of thinking in humans and LRMs by comparing human reaction times with model reasoning traces across a range of reasoning tasks. We isolate this alignment by turning to abductive reasoning: unlike deductive tasks, its difficulty cannot be inferred from formal structure and offers no shortcuts a model could exploit to mimic effort without genuine search, providing firmer ground for empirical claims of shared effort. We find further evidence of alignment between LRM and human reasoning effort, as well as evidence that models and humans tend to make similar errors. Finally, we show that decoding methods that let models explore multiple reasoning paths increase alignment in reasoning cost between humans and LRMs across the three models tested.

推理模型类比推理认知对齐思维成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。