arXiv:2602.10329cs.CLcs.AI2026-02被引 1

模型推理时会自动根据任务复杂度调整计算策略,像人一样理性分配资源。

Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality

  • 通过任务复杂度变化观察模型推理策略转变
  • 无显式成本奖励下仍表现出适应性推理能力
  • 适合关注大模型推理机制与认知类比的研究者

人类推理受资源理性驱动——在约束条件下优化表现。近期推理时扩展计算(inference-time scaling)成为提升大语言模型推理性能的有效方法:指令微调模型(IT)在推理时生成长链条推理,而大推理模型(LRM)通过强化学习训练以发现最优推理路径。但目前尚不清楚这种扩展是否能在没有显式计算成本奖励的情况下催生资源理性。我们设计了一种变量归因任务,要求模型根据候选变量、输入输出样例和预定义逻辑函数推断决定结果的变量。通过调节候选变量数量和样例数系统地改变任务复杂度。结果显示,两类模型均随复杂度上升从暴力枚举转向分析性策略;其中IT模型在异或(XOR)和同或(XNOR)函数上性能下降,而LRM保持稳健。这表明即使无显式成本奖励,模型也能根据任务复杂度动态调整推理行为,为推理时扩展本身可催生资源理性提供了有力证据。

原文摘要 · Abstract (English)

Human reasoning is shaped by resource rationality -- optimizing performance under constraints. Recently, inference-time scaling has emerged as a powerful paradigm to improve the reasoning performance of Large Language Models by expanding test-time computation. Specifically, instruction-tuned (IT) models explicitly generate long reasoning steps during inference, whereas Large Reasoning Models (LRMs) are trained by reinforcement learning to discover reasoning paths that maximize accuracy. However, it remains unclear whether resource-rationality can emerge from such scaling without explicit reward related to computational costs. We introduce a Variable Attribution Task in which models infer which variables determine outcomes given candidate variables, input-output trials, and predefined logical functions. By varying the number of candidate variables and trials, we systematically manipulate task complexity. Both models exhibit a transition from brute-force to analytic strategies as complexity increases. IT models degrade on XOR and XNOR functions, whereas LRMs remain robust. These findings suggest that models can adjust their reasoning behavior in response to task complexity, even without explicit cost-based reward. It provides compelling evidence that resource rationality is an emergent property of inference-time scaling itself.

推理机制资源理性大模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。