测试大模型能否模拟人类风险偏好,发现模型普遍更保守,中文表现偏差更大。
Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study
- 用四城真实问卷数据,让大模型预测人们在彩票任务中的选择。
- 两款模型都比真人更避险,o1-mini与真人更接近,中文任务偏差更大。
- 提示语用中文时模型表现更差,反映语言文化对模拟的影响。
大语言模型在对话系统、内容生成和领域咨询等任务中取得显著进展,但其在模拟复杂决策行为(如风险决策)时的可靠性引发关注。本研究评估大模型在风险决策场景中的表现,对比了来自悉尼、达卡、香港和南京的交通意愿调查数据中的人类实际选择与模型生成结果。向ChatGPT 4o和ChatGPT o1-mini提供人口统计信息,要求其预测个体选择。采用常相对风险规避(CRRA)框架分析风险偏好。结果显示,两款模型均表现出比人类更强烈的避险倾向,其中o1-mini与真实人类选择更为接近。对南京和香港的多语言数据进一步分析表明,中文提示下的模型预测偏差大于英文提示,说明提示语言可能影响模型在跨文化情境下的模拟性能。研究揭示了大模型在复现人类风险行为方面的潜力与局限性,尤其是在语言和文化背景差异下的表现。
原文摘要 · Abstract (English)
Large language models (LLMs) have made significant strides, extending their applications to dialogue systems, automated content creation, and domain-specific advisory tasks. However, as their use grows, concerns have emerged regarding their reliability in simulating complex decision-making behavior, such as risky decision-making, where a single choice can lead to multiple outcomes. This study investigates the ability of LLMs to simulate risky decision-making scenarios. We compare model-generated decisions with actual human responses in a series of lottery-based tasks, using transportation stated preference survey data from participants in Sydney, Dhaka, Hong Kong, and Nanjing. Demographic inputs were provided to two LLMs -- ChatGPT 4o and ChatGPT o1-mini -- which were tasked with predicting individual choices. Risk preferences were analyzed using the Constant Relative Risk Aversion (CRRA) framework. Results show that both models exhibit more risk-averse behavior than human participants, with o1-mini aligning more closely with observed human decisions. Further analysis of multilingual data from Nanjing and Hong Kong indicates that model predictions in Chinese deviate more from actual responses compared to English, suggesting that prompt language may influence simulation performance. These findings highlight both the promise and the current limitations of LLMs in replicating human-like risk behavior, particularly in linguistic and cultural settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。