arXiv:2502.12470cs.CL2025-02被引 23

让大模型像人一样灵活切换直觉与分析思维,提升应对不同任务的能力。

Reasoning on a Spectrum: Aligning LLMs to System 1 and System 2 Thinking

  • 通过构建双类型答案数据集,让模型对齐直觉(系统1)和分析(系统2)推理。
  • 系统1模型在常识题上更优,系统2在算术符号题上表现更好,存在准确率-效率权衡。
  • 基于生成熵动态融合两类模型,无需再训练即在多数任务上超越单一模型。

大型语言模型虽具备出色推理能力,但高度依赖结构化逐步推理,缺乏人类认知中根据情境灵活切换直觉性(系统1)与分析性(系统2)思维的适应性。本文通过构建包含有效系统1与系统2答案的数据集,评估模型在多个推理基准上的表现。结果表明:系统2对齐模型在算术与符号推理任务中更优,而系统1对齐模型在常识推理任务中表现更佳,二者呈现准确率-效率权衡。通过调节对齐数据比例进行插值,发现准确率随策略连续变化。机制分析显示,系统1模型输出更确定,系统2模型则表现出更高不确定性。在此基础上,仅依据生成熵动态融合两类模型,无需额外训练,即可在几乎所有基准上取得更优性能。该研究挑战了‘逐步推理始终最优’的假设,强调应根据任务需求动态适配推理策略。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit impressive reasoning abilities, yet their reliance on structured step-by-step processing reveals a critical limitation. In contrast, human cognition fluidly adapts between intuitive, heuristic (System 1) and analytical, deliberative (System 2) reasoning depending on the context. This difference between human cognitive flexibility and LLMs' reliance on a single reasoning style raises a critical question: while human fast heuristic reasoning evolved for its efficiency and adaptability, is a uniform reasoning approach truly optimal for LLMs, or does its inflexibility make them brittle and unreliable when faced with tasks demanding more agile, intuitive responses? To answer these questions, we explicitly align LLMs to these reasoning styles by curating a dataset with valid System 1 and System 2 answers, and evaluate their performance across reasoning benchmarks. Our results reveal an accuracy-efficiency trade-off: System 2-aligned models excel in arithmetic and symbolic reasoning, while System 1-aligned models perform better in commonsense reasoning tasks. To analyze the reasoning spectrum, we interpolated between the two extremes by varying the proportion of alignment data, which resulted in a monotonic change in accuracy. A mechanistic analysis of model responses shows that System 1 models employ more definitive outputs, whereas System 2 models demonstrate greater uncertainty. Building on these findings, we further combine System 1- and System 2-aligned models based on the entropy of their generations, without additional training, and obtain a dynamic model that outperforms across nearly all benchmarks. This work challenges the assumption that step-by-step reasoning is always optimal and highlights the need for adapting reasoning strategies based on task demands.

大模型推理认知模拟动态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。