arXiv:2504.10368cs.CLcs.AI2025-04中稿 · IJCAI被引 14

探究大模型的直觉式快速推理能力,发现其效率与准确率仍有提升空间。

Exploring the System 1 Thinking Capability of Large Reasoning Models

  • 构建S1-Bench多领域多语言基准,测试模型直觉反应能力
  • 28个大模型在简单问题上表现欠佳,效率与准确率双低
  • 模型早期已感知难度但信心下降,隐状态中编码了难度信息

本文探索大推理模型(LRMs)的系统1思维能力,即以极少词元高效响应的直觉化推理能力。尽管现有LRMs依赖长链推理并在复杂任务中表现出色,其系统1思维能力仍鲜有研究。该能力对反映模型的难度感知与推理效率至关重要,直接影响实际应用效果。为此,我们提出S1-Bench,一个涵盖多领域、多语言的基准,包含模型-简单系统1问题。对28个LRMs的评估显示,它们在系统1任务上存在显著的不准确与低效问题。进一步分析表明,现有高效推理方法要么在简单问题上泛化能力差,要么为追求效率牺牲性能。此外,我们发现模型在早期已具备难度感知,伴随置信度下降,并且问题难度被隐式编码于隐藏状态中。

原文摘要 · Abstract (English)

This paper explores the system 1 thinking capability of Large Reasoning Models (LRMs), the intuitive ability to respond efficiently with minimal token usage. While existing LRMs rely on long-chain reasoning and excel at complex tasks, their system 1 thinking ability remains largely underexplored. This capability is essential as it reflects models' difficulty awareness and reasoning efficiency, both critical for real-world applications. We propose S1-Bench, a multi-domain, multilingual benchmark comprising model-simple system 1 questions. Our investigation of 28 LRMs reveals under-accuracy and inefficiency on system 1 problems. We find existing efficient reasoning methods either generalize poorly to simple questions or sacrifice performance for efficiency. Further exploration uncovers LRMs' early difficulty awareness accompanied by lower confidence, and shows that problem difficulty is implicitly encoded in hidden states.

大模型推理系统1思维效率评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。