arXiv:2505.09901cs.LGcs.AI2025-05被引 2

对比大模型与人类在决策中的探索与利用策略,发现思考能力让大模型更像人。

Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments

  • 用经典老虎机实验比较大模型、人类和算法的探索行为。
  • 有思考能力的大模型在简单环境里探索方式接近人类,复杂环境仍落后。
  • 适合研究智能体行为模拟或自动化决策的学者参考。

大型语言模型(LLMs)越来越多地被用于在复杂序列决策场景中模拟或自动化人类行为。一个核心问题是:LLMs 是否表现出与人类相似的决策行为,并能达到相当(或更优)的性能?本文聚焦于探索-利用(E&E)权衡这一动态决策中的基本问题。我们采用认知科学和精神病学文献中经典的多臂老虎机(MAB)实验,对 LLMs、人类及 MAB 算法的 E&E 策略进行对比研究。通过可解释的选择模型捕捉各智能体的 E&E 策略,并探究通过提示策略和思维模型启用思考过程如何影响 LLM 的决策。结果发现,赋予思考能力使 LLM 的行为更趋近于人类,表现为随机探索与定向探索的混合。在简单静态环境下,思考增强的 LLM 与人类在随机和定向探索水平上表现相似;但在更复杂的非平稳环境中,尽管部分情景下遗憾值相近,LLMs 仍难以匹敌人类的适应性,尤其是在有效定向探索方面。研究揭示了 LLM 作为人类行为模拟器和自动化决策工具的潜力与局限,并指明改进方向。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A natural question is then whether LLMs exhibit similar decision-making behavior to humans, and can achieve comparable (or superior) performance. In this work, we focus on the exploration-exploitation (E&E) tradeoff, a fundamental aspect of dynamic decision-making under uncertainty. We employ canonical multi-armed bandit (MAB) experiments introduced in the cognitive science and psychiatry literature to conduct a comparative study of the E&E strategies of LLMs, humans, and MAB algorithms. We use interpretable choice models to capture the E&E strategies of the agents and investigate how enabling thinking traces, through both prompting strategies and thinking models, shapes LLM decision-making. We find that enabling thinking in LLMs shifts their behavior toward more human-like behavior, characterized by a mix of random and directed exploration. In a simple stationary setting, thinking-enabled LLMs exhibit similar levels of random and directed exploration compared to humans. However, in more complex, non-stationary environments, LLMs struggle to match human adaptability, particularly in effective directed exploration, despite achieving similar regret in certain scenarios. Our findings highlight both the promise and limits of LLMs as simulators of human behavior and tools for automated decision-making and point to potential areas for improvement.

决策模型大模型行为探索利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。