arXiv:2501.18009cs.AIq-bio.NC2025-01NeurIPS被引 10

大模型探索能力弱,因思考太快导致决策过早。

Large Language Models Think Too Fast To Explore Effectively

  • 用拼图游戏测试大模型探索能力,发现多数表现不如人类。
  • 传统模型推理快但粗糙,仅依赖不确定性驱动策略。
  • 深度思考型模型更像人,适合需要创造性探索的任务。

大型语言模型(LLMs)展现出多种智能能力,但对其探索能力的关注仍有限——这是自然与人工系统发现新信息、适应新环境的关键能力。本研究以《小炼金术2》为范式,考察大模型在开放式任务中是否能超越人类的探索表现。结果表明,除o1模型外,大多数传统大模型表现逊于人类,主要依赖不确定性驱动策略,而人类则平衡不确定性与自主性。传统推理型模型如GPT-4o表现出显著更快、更简略的推理过程,限制了探索效果。相比之下,DeepSeek推理模型展现出持续迭代的思考模式,反复分析组合与过往尝试,更接近人类探索策略。通过稀疏自编码器(SAE)的表征分析发现,不确定性与选择在早期变换器层被表示,而自主性值则在后期处理,导致模型思考过快、决策过早,阻碍有效探索。该研究揭示了大模型探索的局限性,并指明提升其适应性的方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged with many intellectual capacities. While numerous benchmarks assess their intelligence, limited attention has been given to their ability to explore--an essential capacity for discovering new information and adapting to novel environments in both natural and artificial systems. The extent to which LLMs can effectively explore, particularly in open-ended tasks, remains unclear. This study investigates whether LLMs can surpass humans in exploration during an open-ended task, using Little Alchemy 2 as a paradigm, where agents combine elements to discover new ones. Results show most LLMs underperform compared to humans, except for the o1 model, with traditional LLMs relying primarily on uncertainty-driven strategies, unlike humans who balance uncertainty and empowerment. Results indicate that traditional reasoning-focused LLMs, such as GPT-4o, exhibit a significantly faster and less detailed reasoning process, limiting their exploratory performance. In contrast, the DeepSeek reasoning model demonstrates prolonged, iterative thought processes marked by repetitive analysis of combinations and past trials, reflecting a more thorough and human-like exploration strategy. Representational analysis of the models with Sparse Autoencoders (SAE) revealed that uncertainty and choices are represented at earlier transformer blocks, while empowerment values are processed later, causing LLMs to think too fast and make premature decisions, hindering effective exploration. These findings shed light on the limitations of LLM exploration and suggest directions for improving their adaptability.

探索能力大模型推理机制认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。