LLM决策接近最优,但学习方式与人类完全不同。
Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior
- 用心理学实验任务测试5个主流LLM的决策能力。
- 在不确定性、风险和变通性上,多数LLM表现优于人类。
- 适合关注AI决策机制或人机协同的研究者阅读。
人类决策是社会与文明的基础,而未来大量决策将交由人工智能处理。大语言模型(LLMs)已改变人工智能辅助决策的形态与范围,但其决策学习过程与人类的差异仍不清晰。本研究通过三个经典心理学实验任务,评估了五个主流LLM在真实决策中的三大核心维度:不确定性、风险与认知转换。我们以360名新招募的人类参与者为基准进行对比。结果表明,所有任务中LLM的表现常优于人类,接近近似最优水平。然而,其决策机制与人类存在根本差异。一方面,这证明了LLM具备管理不确定性、校准风险和适应变化的能力;另一方面,这种差异凸显了将其直接替代人类判断的风险,亟需进一步研究。
原文摘要 · Abstract (English)
Human decision-making belongs to the foundation of our society and civilization, but we are on the verge of a future where much of it will be delegated to artificial intelligence. The arrival of Large Language Models (LLMs) has transformed the nature and scope of AI-supported decision-making; however, the process by which they learn to make decisions, compared to humans, remains poorly understood. In this study, we examined the decision-making behavior of five leading LLMs across three core dimensions of real-world decision-making: uncertainty, risk, and set-shifting. Using three well-established experimental psychology tasks designed to probe these dimensions, we benchmarked LLMs against 360 newly recruited human participants. Across all tasks, LLMs often outperformed humans, approaching near-optimal performance. Moreover, the processes underlying their decisions diverged fundamentally from those of humans. On the one hand, our finding demonstrates the ability of LLMs to manage uncertainty, calibrate risk, and adapt to changes. On the other hand, this disparity highlights the risks of relying on them as substitutes for human judgment, calling for further inquiry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。