测试大模型在金融模拟中的二元决策偏见,发现不同模型响应差异巨大。
Evaluating Binary Decision Biases in Large Language Models: Implications for Fair Agent-Based Financial Simulations
- 用单次和少样本调用测试GPT模型决策分布
- GPT-4o-Mini响应率32-43%,而GPT-4-0125-preview达98-99%同意
- 采样方式与模型版本显著影响结果,适合金融仿真研究者参考
大型语言模型(LLMs)正被用于构建类人决策的代理型金融市场模型(ABMs)。随着模型能力增强与普及,研究者可将个体LLM决策纳入ABM环境。然而,集成可能引入内在偏见,需谨慎评估。本文测试三种先进GPT模型的偏见,采用单次与少样本API调用两种方法。结果显示,特定模型及其子版本间输出分布存在显著差异:GPT-4o-Mini-2024-07-18表现较好(32-43%同意),而GPT-4-0125-preview呈现极端偏倚(98-99%同意)。采样方法与模型子版本显著影响结果——独立调用与批量调用产生不同分布。目前无任何GPT模型能在单次测试中同时满足均匀分布与马尔可夫性,但少样本采样在特定条件下可接近均匀分布。我们分析了Temperature参数,提供定义与对比结果。进一步将结果与真实随机二元序列比较,检验人类常见的负近期性偏见,发现LLMs在此方面表现混合,部分情境下能超越人类。这些发现强调在金融市场及其他领域进行严谨的LLM集成的重要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being used to simulate human-like decision making in agent-based financial market models (ABMs). As models become more powerful and accessible, researchers can now incorporate individual LLM decisions into ABM environments. However, integration may introduce inherent biases that need careful evaluation. In this paper we test three state-of-the-art GPT models for bias using two model sampling approaches: one-shot and few-shot API queries. We observe significant variations in distributions of outputs between specific models, and model sub versions, with GPT-4o-Mini-2024-07-18 showing notably better performance (32-43% yes responses) compared to GPT-4-0125-preview's extreme bias (98-99% yes responses). We show that sampling methods and model sub-versions significantly impact results: repeated independent API calls produce different distributions compared to batch sampling within a single call. While no current GPT model can simultaneously achieve a uniform distribution and Markovian properties in one-shot testing, few-shot sampling can approach uniform distributions under certain conditions. We explore the Temperature parameter, providing a definition and comparative results. We further compare our results to true random binary series and test specifically for the common human bias of Negative Recency - finding LLMs have a mixed ability to 'beat' humans in this one regard. These findings emphasise the critical importance of careful LLM integration into ABMs for financial markets and more broadly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。