arXiv:2504.10636econ.GNcs.AI2025-04被引 1

对比人类与ChatGPT在贝叶斯决策中的表现,发现新版AI已接近完美。

Who is More Bayesian: Humans or ChatGPT?

  • 用实验室数据对比人类与ChatGPT在二分类任务中的决策行为。
  • 早期ChatGPT(3.5)表现低于人类,最新版(4o)几乎完全符合贝叶斯规则。
  • 适合关注AI认知能力演进与人类决策偏差的研究者阅读。

我们比较了人类与人工智能(AI)决策者在简单二分类任务中的表现,其中最优决策规则由贝叶斯法则给出。重新分析了El-Gamal和Grether、Holt和Smith实验室实验中收集的人类选择数据。结果表明,尽管总体上贝叶斯法则仍是预测人类选择的最佳模型,但受试者存在异质性,相当一部分人因判断偏差做出次优决策,如‘代表性直觉’(过度重视样本证据而轻视先验)和‘保守主义’(过度重视先验而轻视样本)。我们对比了近期大型语言模型(LLMs)如多个版本的ChatGPT的表现。这些通用生成式对话机器人并未专门训练于狭窄决策任务,而是以网络文本为语料进行‘语言预测’训练。结果显示,ChatGPT同样受制于导致次优决策的偏差。然而我们记录到其性能迅速提升:早期版本(ChatGPT 3.5)表现低于人类,而最新版本(ChatGPT 4o)已达到超人类水平,近乎完美地遵循贝叶斯分类。

原文摘要 · Abstract (English)

We compare the performance of human and artificially intelligent (AI) decision makers in simple binary classification tasks where the optimal decision rule is given by Bayes Rule. We reanalyze choices of human subjects gathered from laboratory experiments conducted by El-Gamal and Grether and Holt and Smith. We confirm that while overall, Bayes Rule represents the single best model for predicting human choices, subjects are heterogeneous and a significant share of them make suboptimal choices that reflect judgement biases described by Kahneman and Tversky that include the ``representativeness heuristic'' (excessive weight on the evidence from the sample relative to the prior) and ``conservatism'' (excessive weight on the prior relative to the sample). We compare the performance of AI subjects gathered from recent versions of large language models (LLMs) including several versions of ChatGPT. These general-purpose generative AI chatbots are not specifically trained to do well in narrow decision making tasks, but are trained instead as ``language predictors'' using a large corpus of textual data from the web. We show that ChatGPT is also subject to biases that result in suboptimal decisions. However we document a rapid evolution in the performance of ChatGPT from sub-human performance for early versions (ChatGPT 3.5) to superhuman and nearly perfect Bayesian classifications in the latest versions (ChatGPT 4o).

贝叶斯决策大模型行为认知偏差AI进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。