发现大模型在经济决策中存在系统性偏见,且可通过提示纠正。
Behavioral Economics of AI: LLM Biases and Corrections
- 用人类实验框架测试大模型行为,对比不同规模版本表现。
- 模型越大越像人,在偏好任务中表现更类人,信念任务中更理性。
- 引导提示可有效降低偏见,适合研究AI决策机制的学者。
生成式AI模型,尤其是大语言模型(LLMs),在经济与金融决策中是否表现出系统性行为偏见?若存在,又该如何缓解?基于认知心理学与实验经济学文献,我们开展了迄今最全面的实验——这些实验原本旨在记录人类偏见——应用于主流大模型家族,覆盖不同模型版本与规模。结果揭示了大模型行为的系统性模式:在偏好类任务中,随着模型规模或先进性的提升,其响应愈发趋近人类;而在信念类任务中,先进的大规模模型常产生理性回应。通过提示模型做出理性决策,可显著减少偏见。
原文摘要 · Abstract (English)
Do generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions? If so, how can these biases be mitigated? Drawing on the cognitive psychology and experimental economics literatures, we conduct the most comprehensive set of experiments to date$-$originally designed to document human biases$-$on prominent LLM families across model versions and scales. We document systematic patterns in LLM behavior. In preference-based tasks, responses become more human-like as models become more advanced or larger, while in belief-based tasks, advanced large-scale models frequently generate rational responses. Prompting LLMs to make rational decisions reduces biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。