测试五大AI聊天机器人在行为经济学游戏中的决策表现
How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games
- 用行为经济学游戏评测5大主流AI聊天机器人的决策策略
- 发现不同模型在博弈中表现出明显的行为差异
- 适合关注AI决策可靠性与应用风险的研究者参考
大型语言模型(LLMs)在各类应用中的部署,要求深入理解其决策策略与行为模式。作为近期行为图灵测试研究的补充,本文对五种领先的基于LLM的聊天机器人家族在一系列行为经济学游戏中的表现进行了全面分析。通过基准测试,我们旨在揭示并记录这些AI在多种情境下的共性与差异行为模式。研究结果提供了各模型战略偏好的重要洞见,凸显了其在关键决策角色中部署时可能带来的影响。
原文摘要 · Abstract (English)
The deployment of large language models (LLMs) in diverse applications requires a thorough understanding of their decision-making strategies and behavioral patterns. As a supplement to a recent study on the behavioral Turing test, this paper presents a comprehensive analysis of five leading LLM-based chatbot families as they navigate a series of behavioral economics games. By benchmarking these AI chatbots, we aim to uncover and document both common and distinct behavioral patterns across a range of scenarios. The findings provide valuable insights into the strategic preferences of each LLM, highlighting potential implications for their deployment in critical decision-making roles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。