arXiv:2506.22496cs.CYcs.AI2025-06被引 1

用行为经济学方法纠正大模型的赌博式冒险行为

Mitigating Gambling-Like Risk-Taking Behaviors in Large Language Models: A Behavioral Economics Approach to AI Safety

  • 引入风险校准训练与损失厌恶机制,抑制模型冲动决策
  • 实测降低18.7%过度自信、24.3%亏损追投倾向
  • 适合关注AI安全与可信决策的研究者

大型语言模型表现出类似赌博心理的系统性风险偏好,包括过度自信、亏损追投及概率误判。基于行为经济学与前景理论,我们识别并形式化了这些“类赌博”模式:模型为高回报输出牺牲准确性,出错后风险上升,且对不确定性的判断持续失准。提出风险感知响应生成(RARG)框架,融合赌博研究洞见,通过风险校准训练、损失厌恶机制与不确定性感知决策缓解上述偏差。引入基于经典赌博心理学实验的新型评估范式,如适配后的爱荷华赌牌任务与概率学习测试。实验表明,模型的类赌博行为显著降低:过度自信下降18.7%,亏损追投减少24.3%,风险校准能力在多场景下提升。本工作首次建立系统性框架,用于理解与缓解人工智能中的赌博心理模式。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit systematic risk-taking behaviors analogous to those observed in gambling psychology, including overconfidence bias, loss-chasing tendencies, and probability misjudgment. Drawing from behavioral economics and prospect theory, we identify and formalize these "gambling-like" patterns where models sacrifice accuracy for high-reward outputs, exhibit escalating risk-taking after errors, and systematically miscalibrate uncertainty. We propose the Risk-Aware Response Generation (RARG) framework, incorporating insights from gambling research to address these behavioral biases through risk-calibrated training, loss-aversion mechanisms, and uncertainty-aware decision making. Our approach introduces novel evaluation paradigms based on established gambling psychology experiments, including AI adaptations of the Iowa Gambling Task and probability learning assessments. Experimental results demonstrate measurable reductions in gambling-like behaviors: 18.7\% decrease in overconfidence bias, 24.3\% reduction in loss-chasing tendencies, and improved risk calibration across diverse scenarios. This work establishes the first systematic framework for understanding and mitigating gambling psychology patterns in AI systems.

AI安全行为经济学大模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。